Population Fidelity: Evaluating Population Representativeness in LLMs

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses representational biases in large language models simulating human attitudes, specifically the compression of opinion ranges and misrepresentation of subgroups. It introduces the concept of demographic fidelity to distinguish aggregate alignment from the representation of internal variation, and establishes a multidimensional statistical evaluation framework encompassing group-level accuracy, inter-group variance, and structural properties. Through large-scale survey replication and comparative experiments, this work systematically analyzes the performance of various models and cultural fine-tuning strategies. The findings reveal that poor representations stem from erroneous allocation of inter-group variance and expose the limitations of cultural fine-tuning in preserving population diversity. By providing reusable code and benchmarks, this research demonstrates that comprehensive demographic simulation requires the simultaneous reproduction of multiple attitudinal characteristics.
📝 Abstract
Large language models (LLMs) show considerable potential in simulating human attitudes and preferences. Prior work finds that LLM-generated responses can compress the range of attitudes found within populations and misrepresent particular subgroups in ways that vary across models and topics. We introduce Population Fidelity, an evaluation framework that distinguishes key conditions required for a set of LLM-generated responses to represent a population. It incorporates three dimensions: group-level accuracy, the amount of between-group variation, and the structure of that variation. We demonstrate the framework's utility in two ways. First, we reproduce a prior study of "machine bias" in LLM survey responses and apply the framework to its models and more recent ones, showing that poor representation reflects not only insufficient between-group variation but also variation assigned to the wrong groups. Second, we evaluate one proposed approach to improving models' population representativeness: cultural fine-tuning. We find that cultural fine-tuning can improve alignment with the survey center without improving the representation of within-population differences, a distinction that measures of aggregate agreement do not capture. We argue that representing a population requires models to reproduce several features of human attitudinal variation simultaneously. Our framework organizes these features and provides reusable code, data, and trained models for evaluating population fidelity across substantive domains and assessing proposed alignment methods.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Population Fidelity
Population Representativeness
Human Attitudes Simulation
Machine Bias
Innovation

Methods, ideas, or system contributions that make the work stand out.

Population Fidelity
Large Language Models
Evaluation Framework
Cultural Fine-tuning
Group-level Variation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
N
Neemias B. da Silva
1 University of Toronto, Canada; 2 Federal University of Technology – Paraná, Brazil
M
Martin Lukk
1 University of Toronto, Canada
A
Ali Sutani
1 University of Toronto, Canada
A
Abhishek Moturu
1 University of Toronto, Canada
H
Harris Yang
1 University of Toronto, Canada
D
Daniel Silver
1 University of Toronto, Canada
Matt Ratto
Matt Ratto
Professor, Faculty of Information, University of Toronto
philosophy of technologysocial studies of science and technologymaterialitysocial theoryinformation
T
Thiago H. Silva
1 University of Toronto, Canada; 2 Federal University of Technology – Paraná, Brazil