Can Open-Weight Large Language Models (LLMs) Simulate Human Survey Populations? A Cross-Instrument Calibration Study

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the uncertain efficacy of open-source large language models (LLMs) in generating synthetic respondent data, particularly regarding their capacity to preserve authentic human psychometric structures. Focusing on open-weight models, this work proposes a distance-controlled cross-instrument calibration paradigm. Through personality conditioning techniques, conditioned responses from three LLM families are systematically benchmarked against large-scale human panel data. The findings reveal a non-monotonic evolution of model capabilities across versions. Notably, simulations achieve correlations of 0.70–0.73 with actual human responses, demonstrating that unfinetuned open-source models can reproduce salient human psychological structures. However, the validity of such synthetic data remains contingent upon version-specific verification, underscoring the necessity of rigorous evaluation when deploying LLMs for psychometric simulation.
📝 Abstract
Large language models (LLMs) are increasingly used to generate synthetic survey respondents and digital twins of real people, but whether their output preserves real human statistical structure, rather than surface plausibility, remains unresolved, and most existing evidence comes from proprietary models rather than open-weight ones. We evaluate three open-weight LLM families on a cross-instrument calibration task: conditioning personas on real respondents'verbatim answers to one psychometric instrument and measuring them on a second, construct-distance-controlled instrument, checked against a 2,058-person human panel. Across a 139-pair grid, the simulated cross-instrument correlation tracks the real human correlation at r = 0.70 - 0.73 in every model, driven mainly by correct sign rather than precise magnitude and concentrated in pairs of moderate construct distance. A correlation of this magnitude, obtained from untuned open-weight models conditioned only on individual-level survey data, is a substantively encouraging result for LLM-based behavioral simulation and digital-twin applications: specific model families and releases already reproduce a meaningful share of real human cross-instrument structure without any fine-tuning. This capability does not, however, improve monotonically across model releases: on a matched panel, the newest of three tested Llama releases performs worst on two of three headline metrics, so realizing its promise in practice requires release-specific, distance-aware verification rather than a one-time benchmark.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Synthetic Survey Respondents
Digital Twins
Cross-Instrument Calibration
Open-Weight Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Open-Weight LLMs
Cross-Instrument Calibration
Digital Twins
Synthetic Survey Respondents
Behavioral Simulation
🔎 Similar Papers
No similar papers found.