Sahte Kimlik > Makaleler > Where Our Taiwan Name Data Comes From

Bu makale henüz Türkçe diline çevrilmedi — English dilindeki orijinali okuyorsunuz. Şu dillerde de mevcut:Deutsch, English, Українська

Where Our Taiwan Name Data Comes From

Every weight in our Taiwanese surname dataset traces back to a named row in a government registry with a bearer count next to it. Nothing in the head is modelled, interpolated, or borrowed from a neighbouring country. That is not true of any other locale we maintain, and it is worth explaining exactly what it buys — and what it still does not.

This page documents the registry we used, the line between what we measured and what we estimated, the traps specific to Taiwanese onomastics, and the things we have not fixed.

The registry we used

FieldValue
Dataset姓氏排名與人數按三階段年齡分 (Surname rank and population by three age bands)
Publisher內政部戶政司 — Department of Household Registration, Ministry of the Interior
Portalhttps://data.gov.tw/dataset/126774
Reference date30 June 2023 (民國112年6月30日)
Depth2,731 distinct surnames, each split across 3 age bands (0–14, 15–64, 65+)
Fieldsstatistic year, rank, surname, age band, population count
LicenceGovernment Open Data Licence v1
Population base23,373,283

The companion narrative report — 《全國姓名統計分析》, 8th edition, published October 2023 — gives the headline aggregates: the top 10 surnames cover 12,339,778 people, or 52.79% of the population, and the top 100 cover 22,582,375, or 96.62%.

The 2,731 vs 1,785 discrepancy

The same ministry says Taiwan has 1,785 surnames: 1,667 single-character surnames covering 23,337,790 people, plus 118 compound surnames covering 27,404 people. Those two figures plus 8,089 people holding indigenous traditional names or transliterated foreign names sum exactly to the national population. So why does the open dataset carry 2,731 rows?

We do not have a documented answer; the ministry publishes no reconciliation. The most plausible reading is that the 1,785 count covers only the 單姓 + 複姓 categories, while the raw dataset lists the residual 8,089-person bucket as individual rows — roughly 946 rare name strings averaging under nine bearers each. That is arithmetically comfortable, but we have not verified it row by row. The honest statement is: the gap is probably the indigenous and transliterated bucket, and we do not know for certain.

What we measured vs what we estimated

Measured. Every surname in the file, and every weight attached to it, comes from the registry. The weight is the registry's bearer count, stored verbatim:

weight = count          # 陳 → 2618994, not a rescaled 255

No power curve, no rank-based fitting, no expert judgement in the ordering. Our top-10 order — 陳, 林, 黃, 張, 李, 王, 吳, 劉, 蔡, 楊 — is the ministry's published order, not our reconstruction of it.

Historical note. Until the format change, weights were rescaled onto an 8-bit scale as round(255 × count / count(陳)), which introduced a rounding floor: a weight of 1 was worth 2,618,994 / 255 = 10,271 bearers, so any surname below that was over-represented. That forced a depth cut-off at roughly 2,000 bearers, 401 entries. Both constraints are gone: the file now carries real counts and ships the full corpus. Numbers quoted from older builds — 401 surnames, weight 255 for 陳 — refer to that superseded format, which is why they do not match anything you can compute from the shipped file today.

Estimated. Nothing in the weights. The only remaining judgement is which rows to exclude as unusable (see below), not how to score the rows we keep.

Not estimated at all: the age split. The registry breaks every surname across three age bands, and we collapse them. Our file carries no age structure.

Pitfalls specific to Taiwan

There is no privacy threshold — so absence really means zero

Most statistical agencies suppress small counts. Taiwan does not. The registry publishes 643 surnames with exactly one bearer and 406 with two, down to Japanese-origin names like 八谷 and 大久保 held by a single person among 23 million.

That changes what the registry can do. In most countries, a surname missing from official data might be rare, might be suppressed, might be a data error — you cannot tell. In Taiwan, missing means zero bearers, full stop.

We used that. Our earlier list contained 33 surnames lifted from the 百家姓, the 11th-century Hundred Family Surnames primer — 亓官, 漆雕, 東郭, 南門, 段干, 百里, 公孫, 萬俟 and others. Every one has zero bearers in Taiwan. They were there because the list had been built from a classical text instead of a living population. They are gone.

But 澹臺 stayed, and that is the whole lesson

澹臺 is an archaic two-character surname from the same primer, sitting in the same semantic neighbourhood as the ghosts, looking identical to them. It was on our deletion list. It is in the registry with two living bearers, so it stayed.

Nothing about its appearance distinguishes it from 東郭. Only the counter does. The same applies to 歐陽 (7,653), 司徒 (478), 上官 (294), 端木 (186), 諸葛 (180), 皇甫 (115), 慕容 (1) and 呼延 (1) — all real, all retained. A registry is not a way to confirm what you already suspected; it is a way to be told you were wrong about a specific row.

This is also why the Taiwanese cleanup does not license the same cleanup elsewhere. Ukraine has no surname registry, so deleting odd-looking Ukrainian surnames would be deletion by vibe. The difference is the counter, not the obviousness of the defect.

異體字: variant graphs are different surnames, and the registry proves it

Taiwanese household registration records the written form, and several surnames exist in more than one legal orthographic variant (異體字): 温/溫, 黄/黃, 鐘/鍾, 凃/涂/塗, 藍/籃, 龎/龐.

The tempting rule is "Taiwan uses traditional characters, so keep 溫 and drop 温." The registry kills that rule with one number: 温 has 42,171 bearers and 溫 has 13,956. The supposedly wrong form is three times more common. Applying the rule would delete 42,000 real citizens and keep 14,000.

These are not mainland simplifications but legal variants in Taiwanese 戶籍; the resemblance to simplified forms is incidental, because simplification often adopted an existing variant. The ratios go both ways: 黃 1,402,808 vs 黄 34,531 and 鍾 152,550 vs 鐘 28,759 favour the traditional form, but 龎 2,099 vs 龐 1,171 is inverted. There is no rule. There is only the table. We keep both forms of every pair.

Separately, 丘 (4,361) and 邱 (343,469) are not variants at all but distinct surnames — 邱 arose from 丘 under the Qing taboo on Confucius's given name 孔丘.

其他 is not a surname

Rank 147 in the registry is 其他 — "other" — with 5,174 people. It is a service category, and any mechanical "take the top N" pass swallows it and produces a dataset in which 5,174 Taiwanese are named "Other." We excluded it explicitly.

The head is not China's head

Anyone importing mainland assumptions gets Taiwan wrong. 陳 leads at 11.21%, 林 is second at 8.33%, and 王 — the most common surname in mainland China — is only rank 6 at 4.09%. These are different populations with different histories, not dialects of one "Chinese" list.

Why we no longer stop at 401 surnames

Depth used to have a measurable price. Because a weight of 1 was worth 10,271 bearers, every entry below that threshold overstated itself. Summing the implied population across the file gives a "claimed coverage" figure, and under the 8-bit scale it was a hard constraint — this is the table that justified the cut-off in the build of 17 July 2026:

Depth thresholdEntriesClaimed coverageReal coverage
≥ 2,000401109.8%99.61%
≥ 1,000414110.3%99.69%
≥ 500440111.5%99.76%
≥ 100579117.6%99.90%
whole registry2,730212.1%100.00%

Loading all 2,730 surnames on that scale would have produced a dataset claiming two Taiwans: the tail's phantom mass, not its length, was the problem.

That constraint no longer exists. With real bearer counts there is no rounding floor and no phantom mass, so there is no reason to truncate. The shipped file carries all 2,707 usable surnames, weights summing to 23,367,536 against a population base of 23,373,283 — a claimed coverage of 99.98%, which is simply the real coverage, because claimed and real are now the same measurement. The table above is kept because older documents quote its figures; it describes a format we no longer ship.

Our regression suite prints 99.9%, not 99.98%, and that is a third denominator rather than a fourth error. ng_regression_tests.php divides by a rounded present-day population estimate of 23,400,000, whereas this page divides by the registry's own base of 23,373,283. The suite's line — 52.81% top-10-in-corpus · 99.9% coverage · 52.73% product — is internally consistent on that base; against the registry base the same identity reads 52.81% × 99.98% = 52.79%, which is the registry's published top-10 share. Neither is wrong; quoting one against the other's denominator is.

The asymmetry that governed the old build was deliberate: we did not add below 2,000 bearers, but we did not delete real entries already present either (慕容 1, 通 1, 澹臺 2). Adding was limited by the scale; deleting something real was never on the table. Now nothing is limited by the scale, and the rule reduces to its second half.

Top 10

RankSurnameBearers (= stored weight)Share of population
12,618,99411.21%
21,947,5208.33%
31,402,8086.00%
41,239,8805.30%
51,199,9205.13%
6955,0354.09%
7935,6844.00%
8737,6073.16%
9684,5312.93%
10617,7992.64%

Every count above is the registry figure, stored verbatim in the shipped file — there is no longer a reconstruction step, and no row needs a ~. Shares are computed against the 23,373,283 population base and may differ by ±0.02 pp from the ministry's own rounding.

52.79% and 52.81% are two different measurements

These two numbers appear throughout our documentation and they are not a discrepancy. They are the same top ten over two different denominators:

FigureDenominatorSurnamesBearers
52.79%The ministry's full registry, as published2,73123,373,283
52.81%Our corpus, after excluding unusable rows2,70723,367,536

The gap is 24 rows and 5,747 bearers — 0.0246% of the population — and we know exactly which rows they are: the service category 其他 (5,174 bearers, not a surname) plus 23 rows in the Unicode Private Use Area, where the ministry encodes rare characters that have no Unicode assignment and which would render as blank boxes for every reader.

Removing a small slice of the tail leaves the top ten untouched while shrinking the denominator, so our share comes out marginally higher. Computed from the shipped weight file: the top ten sum to 12,339,778, which is 52.8074% of our corpus and 52.7944% of the full registry base. The correct phrasing is therefore "the registry gives 52.79%; our corpus, after excluding 24 unusable rows, gives 52.81%" — never "approximately 52.8%", which is how a real discrepancy would hide.

Known limitations

  • The 2,731 / 1,785 gap is unexplained. Our reading is plausible and unverified.
  • 24 registry rows are excluded — the service category 其他 plus 23 Private Use Area rows. That is the whole of what we drop: 5,747 bearers, 0.0246% of the population. The depth truncation that used to remove ~2,330 tail surnames is gone.
  • The tail is now shipped at full precision, and that is a change in kind. Under the 8-bit scale everything between 2,000 and 10,271 bearers collapsed to weight 1 and overstated itself by up to ~5×. Real counts remove the distortion entirely; any document describing that floor is describing the old format.
  • Age structure is discarded. The registry has it; we do not use it. 郭 is the most-aged surname in Taiwan (ageing index 158.64) and our data cannot express that.
  • The top-10 × coverage identity is not verification. Multiplying our in-corpus top-10 share by the corpus's claimed coverage reproduces the population-level top-10 share, and it does so no matter what is in the file: both sides are computed from the same weights, so the identity reduces to (a/b) × (b/c) = a/c, which is true for any inputs. Our regression suite carries this as informational output with a fixed PASS, not as a test. The weights came from this registry, so agreeing with it proves arithmetic, not truth.

Data as of 2026-07-18

Registry reference date: 30 June 2023. Dataset last reviewed and rebuilt: 18 July 2026. Corpus: 2,707 surnames carrying real bearer counts and covering 99.98% of the Taiwanese population. Files for male and female bearers are byte-identical — Taiwanese surnames do not inflect for gender.

Numbers from older builds do not match this page. Documents written before 18 July 2026 describe a corpus of 401 surnames on an 8-bit weight scale, quoting 99.61% coverage, 109.8% claimed coverage and an in-corpus top-10 of 48.14%. Those figures were correct for that format and are wrong for the shipped file. The equivalents today are 2,707 surnames, 99.98% coverage and an in-corpus top-10 of 52.81%. Verified against the shipped weight file on 18 July 2026.

Sources

← Makaleler