Probabilistic Race and Ethnicity Prediction Using Group-Specific Name Lists

πŸ“… 2026-10-05
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenge of accurately estimating racial disparities in the absence of name–race linkage data. To this end, it proposes ellBISG, a novel framework that pioneers the integration of large language models (LLMs) with proximal causal inference. Specifically, the method leverages LLMs to generate group-specific name lists and achieves probabilistic race prediction through vector embeddings combined with Bayesian-improved surname geocoding. Proximal inference is subsequently introduced to correct residual biases, thereby eliminating the reliance on official demographic statistics inherent in traditional approaches. Extensive validation across multiple datasets demonstrates that, even without prior data, the proposed framework yields well-calibrated and precise estimates of both race probabilities and racial disparities. By overcoming longstanding data constraints, this work establishes a new paradigm for algorithmic fairness research.
πŸ“ Abstract
Statistically valid estimation of racial and ethnic disparities often requires inferring the probability that an individual belongs to a particular racial or ethnic group given only their name and geographic location. The standard approach, Bayesian Improved Surname Geocoding (BISG), relies on group population frequencies for each name. Although the U.S. Census Bureau provides such information for common names and a limited set of racial categories, comparable data do not exist for many racial and ethnic groups and are rarely available outside the U.S. We propose the list-powered BISG ($\ell$BISG) method, which can be used to derive calibrated group probabilities from group-specific name lists. These lists may be compiled based on expert knowledge or generated synthetically using large language models (LLMs), and thus may be subject to unknown biases. Representing names as embeddings, we treat list membership as a proxy prediction task and apply a correction based on proximal inference to recover the target group probabilities. We validate the method on U.S. voter files with self-reported race, on the full-count 1900 U.S. Census, and on the Lebanese voter registry. We find that LLM-generated name lists yield accurate and well-calibrated probabilities as well as precise disparity estimates comparable to those obtained using methods that require name-race data. Thus, $\ell$BISG substantially broadens the applicability of probabilistic race and ethnicity prediction to settings where name-race data are unavailable.
Problem

Research questions and friction points this paper is trying to address.

race and ethnicity prediction
probabilistic inference
racial disparities
name-based geocoding
data scarcity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Probabilistic Race Prediction
List-powered BISG
Proximal Inference
Name Embeddings
Large Language Models
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
K
Kyla Chasalow
Department of Statistics, Harvard University
N
Noah Dasanaike
Department of Government, Harvard University
Kosuke Imai
Kosuke Imai
Professor of Government and of Statistics, Harvard University
applied statisticscausal inferencecomputational social sciencequantitative social sciencepolitical methodology