Abstract
<title>Abstract</title> <p>Responsible research assessment rests on correct attribution, yet much practice still attributes works by name string, which fails for common surnames and romanized non-Latin names. We quantify this name-collision burden over the full population of ORCID-bearing authors on one OpenAlex snapshot (March 2026), some 4.6 million researchers, with a transparent self-defined exact-name operator, so the figures are population values, not estimates. We count how many distinct author entities share each researcher’s normalized name under three strategies, across three name-origin strata proxied by affiliation country. On the conservative strategy (full-name matches restricted to ORCID-bearing entities, our primary measure), the median East-Asian name is shared by 3 real, identified people versus a unique 1 for an Anglophone or Other name. On the naive full-name strategy (an upper bound), the median East-Asian researcher matches 9 entities versus 1; under worst-case initial-plus-surname matching, it collides with 4,846 identities versus 69 (Anglophone) and 34 (Other), with 92% of East-Asian names colliding with at least 100. The contrasts carry large rank-biserial effect sizes (up to r = − 0.74); a negative-binomial regression shows that, after adjusting for publication volume, the East-Asian stratum still carries roughly 6.5 to 8.6 times the expected collision count. The burden is large and unequal across name origins, which motivates attributing works by persistent identifier rather than by name; we present and evaluate SigmaCV, an open, FAIR web application that does so. This is a measurement of name ambiguity, an important precondition for inequitable attribution, not a demonstrated assessment harm.</p>