Back to Search View Original Cite This Article

Abstract

<p>Whether a shop, restaurant, or workshop is named after a person (Mario's, Chez Martine, Mehmet Usta) is a basic but largely unmeasured property of the commercial landscape. We present a harmonized, openly documented dataset of 1,052,019 unique (country, business-name) records spanning 193 countries, derived from OpenStreetMap (OSM) and labeled with six fields per name: whether the name is personal (named after a specific human), the namesake string, the inferred gender and name origin of that namesake, a business category, and whether the venue is a chain. Labels were produced by a large language model (DeepSeek V4 Flash) under a single codebook that doubles as the human-coder instrument, applied uniformly across Latin and non-Latin scripts (CJK, Hangul, Arabic, Cyrillic, Devanagari, Thai, Greek). Records were assembled by a documented multi-tier spatial sampling scheme (G20 economies across ~30 cities each; smaller countries whole-country or single-city; coverage &amp;gt;= 100 venues to retain a country) from ~1.51 million venue rows.We release the labeled corpus with country-level aggregates, the codebook, the full annotation prompt, and the human-validation results: a nine-language, three-coder gold standard (majority vote of native-speaker coders recruited on Prolific under a first-language screener, with embedded attention checks) covering the four highest-volume script families (Latin, Cyrillic, Arabic, CJK). Against this gold standard the personal label reaches 89.5% accuracy (Cohen's kappa = 0.79) across 970 coded names, clearing 85% accuracy in every language and script family tested (per-language 85.0% English to 95.0% French). The secondary fields (namesake, gender, name origin, category) are released as exploratory.The resource is designed for re-use: its six fields and per-name granularity support at least 28 pre-specified re-slices spanning naming frequency (by country, trade, chain status, city size), naming form (possessive vs honorific vs bare) by language, namesake demographics and origin, and OSM-tag-derived signals such as brand, website, and accessibility. A companion paper reports the cultural findings; this descriptor documents the data and the reusable LLM-annotation-plus-human-validation methodology so others can audit, extend, and re-slice the corpus.</p>

Show More

Keywords

name namesake whether country fields

Related Articles

PORE

About

Connect