Bu makale henüz Türkçe diline çevrilmedi — English dilindeki orijinali okuyorsunuz. Şu dillerde de mevcut:Deutsch, English, Українська
What We Got Wrong About Name Data
We maintain weighted name corpora for 64 locales — first names, surnames, and the frequency curves that decide how often each one appears. Most public writing about this kind of work describes the finished thing: here is the method, here are the sources, here is the result.
This is the other article. Eight mistakes we actually made, all caught during a single audit on 17 July 2026. Several were caught by reviewers arguing with us, and at least two we defended before conceding.
We are publishing them because the failure modes turned out to be structural rather than careless. Every one came from a number that looked authoritative and turned out to be measuring a different object than we thought. If you work with frequency data of any kind, most of these are waiting for you too.
(A companion piece, Seven Ways Name Data Lies to You, covers the general taxonomy. This one is the confession.)
1. The zombie statistic: we calibrated Vietnam on a 1992 sample
What we believed. Nguyễn is carried by 38% of Vietnamese people. We did not treat this as a claim requiring verification, because it did not feel like a claim — it felt like background knowledge. Wikipedia says it. Press coverage says it. Faker libraries encode it. Consensus that broad reads as evidence.
What is true. It is one number, repeated. Follow the citation chain and it terminates in a single sample by Lê Trung Hoa from 1992 — small, decades old, and reproduced downstream not because it is good but because it is the only published figure anyone can cite.
Three large samples disagree with it while having opposite regional biases:
| Source | n | Nguyễn | Skew |
|---|---|---|---|
| VNTH01 | 1,682,729 | 30.5% | northern (63:37) |
| SG01 (Ho Chi Minh City) | ~241,000 | 31.5% | southern |
| Nguyen Viet Khoa (2024), Genealogy 8(1):16 | 883,835 | 31.57% | peer-reviewed; Hanoi 39.01% vs HCMC 30.61% |
Two samples biased in opposite directions landing within one percentage point of each other is close to decisive. The real figure is approximately 31%.
The damage was not cosmetic. A frequency curve is calibrated against the share of its top entry, so an inflated divisor compresses the whole distribution — every rank in the country, wrong, from one number nobody thought to question.
The lesson. If a proportion is cited by everyone and traces back to one small primary source, that is not corroboration — that is a single point of failure with good SEO.
2. "~10% of Icelanders have family names" — and it nearly broke a correct file
This is the one that stings, because the number was wrong and our data was right and we were preparing to break our data to match the number.
What we believed. Iceland uses patronymics, but roughly 10% of Icelanders carry inherited family names (ættarnöfn) — Blöndal, Briem, Thorlacius, Möller, Zoëga. Our file contained 4.15% of them. That is a factor-of-two shortfall. It looked like an obvious gap to close.
What is true. Hagstofa Íslands (1 January 2017) breaks the population down as:
| Category | Share |
|---|---|
| Patronymics | 82% |
ættarnöfn (Icelandic family names) | 4% |
| Other surnames, mostly of foreign origin | 14% |
The "10%" is not a measurement. It is 4% + 14%, merged — someone summing two rows that mean different things. The 14% are surnames carried by immigrants and their descendants. They are not ættarnöfn, which is a specific legal category: law nr. 54/1925 banned the creation of new ones while leaving existing ones valid and heritable to this day.
Our file was at 4.15% against a true value of 4%. It was accurate, and we were about to "fix" it into a 2.4× overstatement.
The lesson. Verifying a number is not enough. You have to verify what the number counts. "10% have family names" and "4% have ættarnöfn" are both true sentences about Iceland, and only one of them is about the thing in your file.
(A related trap we did avoid: law nr. 54/1925 is the one that banned new family names. 41/1913 is the opposite law — the one that permitted them. The digits are close enough to swap by accident, and the meaning inverts.)
3. Circular verification: one dataset, counted twice
What we believed. Our Hungarian surname file gives ethnonym-derived surnames — Tóth, Horváth, Németh, Oláh and similar — 8.20% of total weight. The academic estimate, from Farkas, is "7–8% of the surname stock." Our number, produced independently, landed one-fifth of a point outside a published scholarly figure.
We wrote this up as independent confirmation. It felt like the strongest evidence in the audit.
What is true. It is not a confirmation of anything. Our weights are calibrated against registry shares — our Tóth sits at 2.192% against a registry value of 2.19%. Farkas takes his numbers from the same registry.
So there are not two measurements agreeing. There is one dataset, counted twice, arriving back where it started. The agreement was guaranteed before we looked: had our number come out at 8.20% against a truth of 3%, the check would have passed identically. It has no power to fail.
This is the most seductive error on the list, because the output is indistinguishable from success. We caught the same pattern a second time, independently, in the same audit.
The lesson. If your figures are calibrated from source X, then checking them against X — or against any work that cites X — is a consistency check, not a verification. Call it that. Real independence looks like item 1: sources with opposing biases converging anyway. Two pointers to one dataset is one dataset.
4. Bearers versus positions: wrong by exactly a factor of two
What we believed. García is carried by 2,915,761 people in Spain. Portugal's top-10 surnames account for 45.7%. Both figures come straight from registries. We used them as targets.
What is true. Spanish and Portuguese people carry two surnames, so the shares across the whole list sum to roughly 200%, not 100 — and there are two different quantities hiding under the word "frequency."
For García, INE's own columns settle it: 1,446,937 as first surname, and 2,915,761 across both positions. Those are not competing estimates but different objects, and which one you want depends on whether your record carries one surname or two.
Portugal shows the same split: top-10 = 45.7% of bearers, but 22.8% of positions.
Confusing them does not produce a rounding error. It produces a target twice as large as reality, and then everything downstream is bent to reach a number that was never achievable. Our own diagnostic tooling fell for it: fed a bearer share, it reported a defect in the Portuguese file that did not exist.
The lesson. Before using any surname frequency from a two-surname country, establish the denominator. If the source does not say and the shares sum toward 200%, you are looking at bearers.
5. Hanja versus hangul: 정 is 鄭 and 丁
What we believed. Korean surnames are a small, well-documented set. Our records are in hangul, hanja share tables are easy to find, and the two obviously describe the same distribution.
What is true. Hangul is phonetic; hanja is not. The syllable 정 corresponds to two distinct surnames — 鄭 and 丁 — with different lineages, collapsed into one written form.
The consequences are not subtle:
| Counted in | Top-10 share |
|---|---|
| hanja | 63.9% |
| hangul | 65.8% |
That 1.9-point gap is not noise. It is the merging. And it moves the ranking, not just the totals: in hangul, 정 = 4.84% (鄭 + 丁 together), which puts it above 최 at 4.7%. Take the hanja table — the one more commonly published — and 鄭 alone looks smaller, the order flips, and you have a file that ranks Korea's surnames wrong in a way no validator will flag.
The lesson. When two writing systems merge distinctions differently, "the frequency of X" is undefined until you say which script you counted in. Pick the one your data is actually written in, and never repair a discrepancy that is just the other script talking.
6. "The country's top 1,500" is not "our list of 1,500"
What we believed. Our French surname file had a compressed weight curve. The reference we reached for: a real top-1,500 of France runs about 50:1 from the commonest surname to the rarest. Our file has around 1,500 entries. So 50:1 is our target.
What is true. The two things share a number and nothing else.
France's actual top-1,500 is a contiguous block: ranks 1 through 1,500, and rank 1,500 is still a reasonably common surname. Our file is a curated list of 1,500 entries scattered across ranks 1 to 126,772 — the very common, plus a deliberate reach into the deep tail where a surname might have a few dozen bearers nationally. Its own real ratio is 3,351:1.
We were about to compress a curve by roughly 70× because two unrelated lists happened to have the same length.
The lesson. Measure the ratio between the top and the rarest entry in your own file, never in an imagined top-N of the country. A curated list and a rank-contiguous list are different animals wearing the same size label.
7. "Strange" is not "invented" — and vice versa
This one is really two mistakes, pointing in opposite directions, made in the same week.
In Ukraine, we were about to delete real people. Our backlog said the Ukrainian surname file was full of artificial entries generated from productive word-formation patterns rather than a registry, and needed cleaning. The suspicion was reasonable — the file's origin story was exactly that.
Then we sampled it. A random draw of n = 25 from the tail returned 0 artificial surnames. Соловʼяненко is an opera singer. Хомчак is a former Commander-in-Chief of the Armed Forces of Ukraine. They sit in the same tail, looking exactly as odd as the entries we wanted to remove. Tested on deliberately suspicious candidates, the criterion "sounds strange" produced a 21% false-positive rate.
The context settles it: Ukraine has 707,685 unique surnames, averaging about 64 bearers each. Rare-and-strange is not an anomaly there — it is the shape of the country.
In Taiwan, the opposite. Our Chinese surname list carried archaic compound surnames straight out of the 百家姓, an 11th-century text — a poem, not a census. Checked against Taiwan's Ministry of the Interior registry, 33 of them had zero bearers. Deleted, correctly.
And then the trap inside the trap: 澹臺 looks identical in kind — archaic, compound, same source text. We put it on the deletion list by appearance. The registry has it with 2 bearers. It stayed.
The lesson. The variable is not how obvious the defect looks — in both countries it looked obvious. The difference between cleaning a database and ruining one is whether a per-name registry exists. Taiwan publishes counts with no privacy floor, so absence genuinely means zero and deletion is a measurement. Ukraine has no such registry at that depth, so deletion collapses into deleting by ear — and one time in five you delete a real one. When you have a counter, trust it over your intuition. When you have none, do not substitute your ear for one.
8. The mirror defect: we fixed one side and made it worse
What we did. Sweden's surname file was missing its migration layer — a real, verifiable gap against the SCB registry, including names that rank inside the country's top 100. We added them. Straightforwardly correct work.
What happened. Output got worse. Records started pairing Swedish first names with newly-added surnames — the classic implausible mix — and it was more noticeable than before we started.
Why. The first names had the same gap, and we had not looked. A probe of 40 markers against the male first-name file found three hits — 0.22% — and the female file had zero. Ahmed did not exist as a first name in our data at all.
So the implausible pairings were not caused by adding surnames. They were caused by those surnames having no matching first names to pair with. Before the fix, both sides were missing and the gap was invisible through symmetry. Fixing one side did not introduce the problem; it revealed it — and revealing it looked exactly like causing it.
(Closed on both sides the same day: the male first-name layer went from 0.22% to 2.98%, the female from 0% to 0.93%.)
The lesson. When two datasets are joined, a defect present in both can be invisible, and a correct fix to one can look like a regression. Fix both sides or neither — and before concluding your fix broke something, check whether it merely exposed something.
The thread running through all eight
Set them side by side and they are the same mistake:
- 38% and 31% are both "the frequency of
Nguyễn" — over different samples. - 10% and 4% are both "Icelanders with family names" — over different categories.
- 8.20% and 7–8% are both "Hungarian ethnonym surnames" — over the same dataset, counted twice.
- 2,915,761 and 1,446,937 are both "
Garcíain Spain" — over positions versus bearers. - 63.9% and 65.8% are both "Korea's top-10" — over different scripts.
- 50:1 and 3,351:1 are both "the ratio of a 1,500-entry list" — over different lists.
- "Strange" means artificial in Taiwan and normal in Ukraine — over different registries.
- 0.22% is both a fix and a defect — depending on the other side of the join.
Not one of these is a wrong number. Every figure in this article is correct about something. They are wrong only about the thing we thought they were about — and the error is silent every time, because a number arrives without its definition attached.
Which is the actual lesson, duller than any of the eight stories: a figure without an explicit denominator, sample, category and scope is not data. It is a numeral. The work is not finding numbers — numbers are abundant and free. The work is establishing what each one counts, and then refusing to compare two of them until you have.
We got all eight of these wrong first and right second. We will get the ninth one wrong too. The only defence we know is to write the definition next to the number, every time, so the next person has something to disagree with.
Data as of 2026-07-17
All figures verified during an audit of weighted name corpora across 63 locales. ⚠ The corpus held 63 locales when these figures were measured; en_IE was added on 18 July 2026 and it holds 64 today. Every mistake described here is one we made and corrected on that date.
- Vietnam: hoten.org VNTH01 (n = 1,682,729) and SG01 (n ≈ 241,000, Ho Chi Minh City), CC BY 4.0 as declared in the page text; Nguyen Viet Khoa (2024), "Vietnamese Surnames," Genealogy 8(1):16, n = 883,835 → 31.57%, Hanoi 39.01% vs HCMC 30.61%. <doi.org/10.3390/genealogy80100…; — Compared against the Lê Trung Hoa 1992 sample. ⚠ The often-quoted size of that sample (n = 1,941) comes from our working notes and we could not confirm it: the secondary literature reproduces its 38.4% without stating sample size or methodology. The conclusion that 38% is a zombie figure rests on the three large samples above, not on that number.
- Iceland: Hagstofa Íslands, 1 January 2017 — 4%
ættarnöfn, 82% patronymics, 14% other (largely of foreign origin). Legal background: law nr. 54/1925 (banned new family names), 41/1913 (permitted them), later 37/1991 and 45/1996. <statice.is/> - Hungary: ethnonym share 8.20% of file weight; registry-calibrated (
Tóth2.192% vs registry 2.19%); Farkas's "7–8% of the surname stock" derives from the same registry. Reported here as a consistency check, not a verification. - Spain: INE, Frecuencias de apellidos, Censo as of 01/01/2025 —
García1,446,937 as first surname; 2,915,761 across both positions. <ine.es/daco/daco42/nombyapel/…; - Portugal: top-10 = 45.7% of bearers / 22.8% of positions.
- Korea: top-10 = 63.9% by hanja, 65.8% by hangul; 정 = 4.84% (鄭 + 丁) vs 최 = 4.7%. Shares computed against ~49.7 million citizens, not total population.
- France: curated file spanning ranks 1–126,772; internal ratio 3,351:1 vs ~50:1 for a rank-contiguous top-1,500.
- Ukraine: reproducible random sample n = 25 (draw 20260717) from the low-weight tail of a 2,209-entry surname file; 0 artificial. Demonstrably artificial entries in the whole file: ~2–10. The Ukrainian corpus has since been rebuilt and now holds 2,562 registry-weighted entries. Country context: 707,685 unique surnames, ~64 bearers each. No per-name registry is published at that depth (
ridni.org,forebears.ioandstats.ridni.orgoffer no downloadable dataset; the literature publishes only the top 100). - Taiwan: Ministry of the Interior surname registry (2,731 surnames with per-name counts) — 33 zero-bearer entries removed;
澹臺retained on 2 bearers. <data.gov.tw/dataset/126774> - Sweden: SCB registry — surnames (411,798 entries, covering 96.1% of the population) and
tilltalsnamnfirst names as of 31 December 2022. Migration layer in male first names 0.22% → 2.98%; female 0% → 0.93%. <scb.se/en/finding-statistics/…;
Where a figure is a target, an estimate, or unconfirmed, it is labelled as such above.