Estimating Nationality from Personal Names in a Multinational Cohort: Performance of NamSor Across Country, Regional, and Onomastic Classifications. | AMiner
Estimating Nationality from Personal Names in a Multinational Cohort: Performance of NamSor Across Country, Regional, and Onomastic Classifications.
BACKGROUND:Nationality, ethnicity, and geographic background are frequently required in medical research but are often unavailable in administrative or registry-based datasets. Name-based inference tools have demonstrated good performance for predicting country of origin, yet their ability to approximate legal nationality remains unclear. OBJECTIVE:To evaluate the performance of NamSor in predicting nationality from personal names in a large multinational cohort and to assess whether aggregation into broader geographic or onomastic regions improves classification accuracy. METHODS:This cross-sectional study included 11,989 marathon participants representing 135 nationalities. Self-reported nationality, as recorded in the official race results, served as the reference standard. NamSor predictions were evaluated at the country level, fine/coarse United Nations (UN) regional levels, and predefined onomastic macro-regions. Performance was assessed using classification accuracy (proportion of correct predictions among classified observations) across probability thresholds. RESULTS:Country-level accuracy was 60.2%. Aggregation improved performance to 69.7% for fine UN regions and 75.2% for coarse UN regions. Coarse onomastic macro-regions achieved the highest accuracy (88.3%). Increasing probability thresholds improved accuracy among classified observations (e.g., 92.5% at ≥0.9 at the country level) but substantially reduced the proportion of observations retained for analysis, with similar trade-offs observed for regional and onomastic classifications. CONCLUSIONS:Name-based inference aligns more closely with linguistic-cultural groupings than with exact legal nationality. While country-level prediction showed substantial misclassification in a highly multinational setting, aggregation into broader regional or onomastic categories markedly improved performance. Broader regional or onomastic classifications may therefore represent a pragmatic alternative when direct nationality data are unavailable.