Sahte Kimlik > Makaleler > Which Countries Share Names? We Measured All 1,953 Pairs — and Argentina's Closest Match Isn't Spain

Bu makale henüz Türkçe diline çevrilmedi — English dilindeki orijinali okuyorsunuz. Şu dillerde de mevcut:Deutsch, English, Українська

Which Countries Share Names? We Measured All 1,953 Pairs — and Argentina's Closest Match Isn't Spain

We took name corpora for all 63 locales the corpus held at the time of the run, compared every one of the 1,953 possible pairs on two similarity metrics, and ranked them. Most of the top of the list is boring in the best way: France and Belgium, Austria and Germany, Peru and Venezuela. History behaved.

Then, at rank 10 on surname similarity — ahead of Spain and Peru, ahead of Germany and Switzerland — sits Argentina and Italy. Two countries that share no border, no language and no colonial relationship. They share 565 surnames, and those surnames carry 26.9% of Argentina's surname mass while being absent from Spain entirely.

That is not a data error. It is 1880–1930, when roughly two million Italians moved to Argentina, arriving in a country of eight million people. The surnames came with them and stayed. Our matrix rediscovered a migration wave from string overlap alone, without being told it happened.

This article walks the whole matrix: what is close, what is far, what surprised us, and — importantly — the two places where the method breaks and produces confident nonsense.

How similarity was measured

Two metrics, deliberately, because they disagree in informative ways.

Jaccard compares the sets. Of all names appearing in either country, what fraction appears in both? It ignores frequency entirely — a name borne by six million people counts the same as one borne by six.

Cosine compares the weighted vectors. Each corpus is normalised to sum to 1, and we take the cosine of the angle between them. A pair scores high only if the shared names carry similar proportions in both countries.

When Jaccard is high and cosine is low, two countries draw from the same pool but use it differently. When cosine is high and Jaccard is low, they overlap on few names but those few are the big ones in both. Both cases are real and both are interesting.

The closest pairs

First names

RankPairJaccardCosineShared names
1Belgium – France0.9070.975759
2Belgium – Switzerland0.8980.955755
3Switzerland – France0.8520.960734
4Canada – Switzerland0.8320.834725
5Peru – Venezuela0.6100.873575
6Moldova – Romania0.5490.594367
7Austria – Germany0.5370.775485
8Italy – Switzerland0.5300.875423
9Jordan – Saudi Arabia0.5160.643370
10Australia – NZ0.5070.873568

(French-language pairs are fr_FR/fr_BE/fr_CH/fr_CA; "Italy – Switzerland" is it_IT/it_CH.)

The French-language cluster dominates so completely that it needs explaining. Belgium and France score 0.907 Jaccard — nine names in ten are shared. No other language family comes close: German-speaking pairs top out at 0.537, Spanish-speaking at 0.610, English-speaking at 0.507.

This is real, but it is real for a specific reason. French given-name repertoires genuinely are more unified across borders than German or English ones, because French naming was centrally regulated for two centuries. From 1803 until 1993, French law restricted legal given names to a state-sanctioned list — saints' names and figures from ancient history — and Belgian and Swiss francophone practice followed the same conventions. Germany never had one list; Bavaria and Prussia named differently. English-speaking countries never had any list at all.

So the French cluster's extraordinary tightness is a fossil of the Napoleonic loi du 11 germinal an XI. Two centuries of one legal repertoire, still visible in a corpus comparison in 2026.

Surnames

RankPairJaccardCosineShared
1Belgium – Switzerland0.4100.327483
2Australia – New Zealand0.3840.8321,240
3Belgium – France0.3770.656638
4Australia – Britain0.3500.804848
5Switzerland – France0.3130.436556
6Britain – New Zealand0.2960.766948
7Portugal – Brazil0.2940.845198
8Belgium – Canada (fr)0.2470.207334
9France – Canada (fr)0.2420.147329
10Argentina – Italy0.2350.171565
11Australia – Canada0.2230.634552
12Canada (en) – Canada (fr)0.2130.535366

Portugal–Brazil is the clearest demonstration of why two metrics matter. Jaccard 0.294 — modest, only 198 shared surnames. Cosine 0.845 — the third-highest in the entire matrix. The two countries share relatively few surname forms, but the ones they share are the dominant ones in both: Silva, Santos, Oliveira, Souza sit at the top of each list. Brazil's repertoire then diverges enormously through Italian, German, Japanese, Lebanese and Indigenous additions that Portugal never received. Same head, different body — exactly the signature you would predict for a colony that stayed linguistically Portuguese while absorbing migration Portugal did not.

Argentina and Italy

Back to rank 10, because it is the finding that made us re-run the script twice.

Argentina and Italy share 565 surnames. Ranked by weight in the Argentine corpus, the top of that shared list reads:

SurnameShare of Argentine surname mass
Rossi0.341%
Russo0.280%
Ferrari0.256%
Esposito0.244%
Bianchi0.207%
Colombo0.207%
Ricci0.207%
Marino0.195%
Conti0.183%
Greco0.183%

The obvious objection is that these might simply be pan-Romance surnames that Spain also has, in which case the Italy link would be spurious. So we removed every surname that also appears in the Spanish corpus. 538 of the 565 survive — and they still carry 26.9% of Argentina's total surname mass.

Roughly a quarter of the weight of the Argentine surname distribution consists of Italian surnames that Spain does not have. Rossi, Ferrari, Esposito, Bianchi did not travel via Madrid. They arrived directly at Buenos Aires between 1880 and 1930, during a migration that at its peak had Italians arriving in Argentina faster than in any other destination including the United States. Argentine demographic scholarship routinely estimates that a majority of Argentines have some Italian ancestry; our surname mass figure is an independent measurement of the same phenomenon from a completely different direction.

The metric that catches this is Jaccard, not cosine. Cosine for the pair is a modest 0.171, because the biggest Argentine surnames are still Spanish — González, Rodríguez, Gómez — and Italy has none of those. The Italian layer sits underneath the Spanish head. Set overlap sees it; weighted overlap partly hides it. If we had run only cosine, we would have missed the single most interesting finding in the matrix.

Direction matters — and usually gets dropped

Similarity is symmetric. Overlap is not, and this is where most name-comparison work goes wrong.

Ask "what share of country A's name mass consists of names that also exist in country B" and you get two different numbers depending on which way you point it:

A → BA's mass found in BB's mass found in AShared names
Japan → Taiwan75.4%8.4%339
Turkey → Germany17.1%2.0%24
Vietnam → Czechia (given)18.1%0.5%21
Portugal → Brazil90.3%73.0%198
Canada (fr) → Canada (en)70.7%34.4%366
India → Britain20.0%5.1%47
Russia → Ukraine0.1%1.8%12

Turkey and Germany. Only 24 surnames are shared — but those 24 are Turkey's biggest: Yılmaz, Kaya, Demir, Çelik, Yıldız, Şahin, Öztürk. They account for 17.1% of Turkish surname mass. From the German side the same 24 names are 2.0% of German mass. Both numbers are correct and they describe the same object: post-1961 Turkish labour migration to West Germany, which put Turkish surnames into German registries at percentage-level frequency without displacing anything. Written as "Turkey and Germany are 17% similar," it would be misleading. Written as "Turkey's most common surnames now carry measurable weight in Germany's," it is right.

Vietnam and Czechia. 21 shared given names — Anh, Minh, Lan, Linh, Mai, Trang — carrying 18.1% of Vietnamese given-name mass and 0.5% of Czech. Czechia hosts one of Europe's largest Vietnamese communities per capita, a legacy of socialist-era labour and study agreements between Czechoslovakia and North Vietnam. The Czech corpus records those names at weight 2 against 96 for Jan. Small in Czechia, enormous in Vietnam, entirely real in both.

Russia and Ukraine. The most striking low number in the matrix. Two adjacent Slavic countries writing in the same alphabet share 12 surnames out of thousands, with Jaccard 0.003. This is not a script problem — both are Cyrillic. It is a genuine structural difference: Russian surnames are overwhelmingly adjectival possessives in -ov/-ev/-in, Ukrainian ones are dominated by the patronymic -enko and the -uk/-yuk family. The suffix systems barely intersect. Ukraine and Russia are further apart on surnames than Britain and India are.

Where the method breaks

Two failure modes, both severe, both easy to miss.

The script barrier

981 of 1,953 pairs share zero given names. 1,287 of 1,953 share zero surnames. Half to two-thirds of our entire matrix is composed of zeros.

Almost none of those zeros mean the naming systems are unrelated. They mean the two corpora are written in different alphabets, and string comparison cannot see through an alphabet.

The proof case is decisive. Serbian appears in our data twice — sr_RS in Cyrillic and sr_Latn_RS in Latin. Same country, same language, same names, and by inspection the two files are the same data in two scripts: identical entry counts (1,143 given names, 958 surnames) and identical weight totals. Their measured similarity:

Metricsr_RS vs sr_Latn_RS
First-name Jaccard0.0000
First-name cosine0.0000
Surname Jaccard0.0000
Surname cosine0.0000
Shared entries0

Perfect dissimilarity between a corpus and a transliteration of itself. Any conclusion drawn from a zero in this matrix without checking the scripts first is worthless — and that includes the entire "most distant pairs" ranking, which is why we do not publish one. The most distant pairs by measurement are Greek and Korean, Hebrew and Bengali, Georgian and Japanese; the measurement is entirely an artefact of the writing systems and tells you nothing about the names.

Seven locales are complete isolates in our matrix, sharing nothing with anyone: Bengali, Greek, Hebrew, Armenian, Georgian, Korean and Nepali. Each is the sole occupant of its script in our corpus.

Han characters are shared; surnames are not

Japan's corpus shares 339 surname entries with Taiwan's, covering 75.4% of Japanese surname mass. Read naively, that says Japan and Taiwan have nearly the same surnames. They do not.

The shared entries are written in Han characters, which both writing systems use. 佐藤 (Satō) genuinely appears in Taiwan's registry — Taiwan's Ministry of the Interior publishes with no suppression threshold, so every surname with at least one bearer is listed, and Taiwan has residents with Japanese surnames from the 1895–1945 colonial period and from later naturalisation. But the weight is:

SurnameShare of TaiwanShare of Japan
佐藤 Satō0.0003%2.609%
鈴木 Suzuki0.0002%2.426%
高橋 Takahashi0.0002%1.905%
陳 Chen11.208%absent

The reverse direction gives 8.4%, and even that overstates it, because it is dominated by single-character surnames like 林 and 王 that are common in both languages for the unrelated reason that both inherited them from Classical Chinese.

The lesson generalises: a shared script inflates overlap, a different script destroys it, and neither has anything to do with whether the names are related. Every cross-script number in this article is stated in one direction with its counterpart shown next to it, for exactly this reason.

What we expected and did not get

Four predictions we made before running the matrix, and how they fared:

PredictionResultWhat actually happened
sr_RS ≈ sr_Latn_RS (near-identical)Failed0.0000 on every metric. Script barrier.
ro_RO ≈ ro_MD (very close)Held0.549 Jaccard on given names — 6th in the matrix.
de_DE ≈ de_AT ≈ de_CHPartialDE–AT 0.537, but DE–CH only 0.373 and AT–CH 0.394. Swiss German naming is further from Germany than expected.
es_* cluster tightlyPartialPeru–Venezuela 0.610, but Spain–Argentina only 0.428 — lower than Australia–New Zealand.

The Spanish result is the second genuine surprise. Spain is not the centre of the Spanish-speaking name cluster. The tightest Hispanic pairs are all American — Peru–Venezuela 0.610, Argentina–Peru 0.528, Argentina–Venezuela 0.475 — while every pair involving Spain scores lower. Latin American countries resemble each other more than any of them resembles Spain, which is what you would expect from four centuries of independent naming drift plus, in Argentina's case, two million Italians.

Conclusion

Measured across 1,953 pairs, name similarity tracks history with unsettling fidelity: Napoleonic name law in the French cluster, Portuguese colonial continuity in the Brazil pairing, Italian mass migration in the Argentine surname layer, Turkish labour migration in the German registry, Cold War study agreements in the Vietnamese names sitting in Czech data.

It also produces confident nonsense wherever alphabets differ, and half our matrix is that nonsense. The single most important operation in this entire analysis was checking whether Serbian resembled itself. It did not — and until you run that check, you cannot tell a real zero from an alphabet.


Methodology and sources

What was computed. For all 63 locales in corpus version share 1.1.7 — the version current at the time of the run; the corpus is now share 1.1.8 — we built two vectors per locale: given names (male and female files merged) and surnames (lastname_male; in most locales the male and female surname files are identical, so merging both would only double every weight). For each of the 1,953 unordered pairs we computed Jaccard on the key sets and cosine on L1-normalised weight vectors, plus directional weighted overlap in both directions. Analysis run 2026-07-18. ⚠ The corpus held 63 locales when these figures were measured; en_IE was added on 18 July 2026 and it holds 64 today. The 1,953 pairs are all unordered pairs of those 63; re-running against 64 would give 2,016.

Limitations, in order of severity.

  1. Script incomparability. String equality is the only matching rule used. No transliteration, no normalisation across alphabets. 981 given-name pairs and 1,287 surname pairs measure as zero, and essentially all of these are script artefacts rather than findings. Demonstrated by the sr_RS / sr_Latn_RS control, which scores 0.0000 against a transliteration of itself. No "most distant pairs" ranking is published in this article, because that ranking would be a ranking of alphabets.
  1. Shared-script inflation. Han-character locales (ja_JP, zh_CN, zh_TW, ko_KR where hanja are used) show elevated overlap that reflects a common character inventory rather than common surnames. Every such figure is reported directionally with both values shown.
  1. Weight scales are not comparable across locales. In the run, 61 of the 63 locales carried relative integers on a 1–255 scale and only zh_TW and vi_VN carried raw head counts. That split has since moved: as of 22 July 2026, 12 of the 64 surname corpora carry registry head counts and 52 are relative. Cosine is computed on L1-normalised vectors, which removes scale but not shape differences arising from differing calibration depth. Jaccard is unaffected, being set-based. Absolute weights are never compared across locales anywhere in this article.
  1. Corpus depth varies from 184 entries (Korean surnames) to 3,830 (Swedish given names). Jaccard is sensitive to this: a deeper corpus on one side mechanically lowers the score by enlarging the union. Pairs of similar depth are more reliable, which is a further reason the French cluster (all four corpora at 798 given names) scores so cleanly.
  1. Corpora are samples, not registries. Except for zh_TW and vi_VN, these are curated frequency-weighted samples built from registry data, not the registries themselves. A name absent from a corpus is not necessarily absent from the country.

Historical claims and where they come from. The name-data measurements above are ours. The historical explanations are not, and are stated as context rather than as findings:

ClaimStatus
565 shared AR–IT surnames; 538 absent from Spain; 26.9% of AR massOur computation, corpus share 1.1.7
~2 million Italians migrated to Argentina 1880–1930Standard migration historiography; context only, not derived from our data
French given names legally restricted 1803–1993Loi du 11 germinal an XI, repealed 1993; context only
Turkish labour migration to West Germany from 19611961 German–Turkish recruitment agreement; context only
Czechoslovak–North Vietnamese labour and study agreementsContext only
Taiwan publishes surnames with no suppression thresholdMinistry of the Interior, 全國姓名統計分析, 2023
Taiwanese residents with Japanese surnames from 1895–1945 and naturalisationInference consistent with the registry data; not independently verified

Where our measurement and the history agree, that is corroboration, not proof — the corpora were built from registry sources that are themselves shaped by these histories. The Argentine–Italian finding is presented as an independent measurement because nothing in our corpus construction pipeline was aware of Italian migration to Argentina; the overlap emerged from string comparison alone.

Related reading: How Migration Plants Surnames covers surname arrival in registries; Migration Layers in Birth Cohorts does the same for given names on a time axis; How to Audit Name Frequency Data covers verification method generally.

← Makaleler