🤖 AI Summary
This study addresses the uncertain efficacy of open-source large language models (LLMs) in generating synthetic respondent data, particularly regarding their capacity to preserve authentic human psychometric structures. Focusing on open-weight models, this work proposes a distance-controlled cross-instrument calibration paradigm. Through personality conditioning techniques, conditioned responses from three LLM families are systematically benchmarked against large-scale human panel data. The findings reveal a non-monotonic evolution of model capabilities across versions. Notably, simulations achieve correlations of 0.70–0.73 with actual human responses, demonstrating that unfinetuned open-source models can reproduce salient human psychological structures. However, the validity of such synthetic data remains contingent upon version-specific verification, underscoring the necessity of rigorous evaluation when deploying LLMs for psychometric simulation.
📝 Abstract
Large language models (LLMs) are increasingly used to generate synthetic survey respondents and digital twins of real people, but whether their output preserves real human statistical structure, rather than surface plausibility, remains unresolved, and most existing evidence comes from proprietary models rather than open-weight ones. We evaluate three open-weight LLM families on a cross-instrument calibration task: conditioning personas on real respondents'verbatim answers to one psychometric instrument and measuring them on a second, construct-distance-controlled instrument, checked against a 2,058-person human panel. Across a 139-pair grid, the simulated cross-instrument correlation tracks the real human correlation at r = 0.70 - 0.73 in every model, driven mainly by correct sign rather than precise magnitude and concentrated in pairs of moderate construct distance. A correlation of this magnitude, obtained from untuned open-weight models conditioned only on individual-level survey data, is a substantively encouraging result for LLM-based behavioral simulation and digital-twin applications: specific model families and releases already reproduce a meaningful share of real human cross-instrument structure without any fine-tuning. This capability does not, however, improve monotonically across model releases: on a matched panel, the newest of three tested Llama releases performs worst on two of three headline metrics, so realizing its promise in practice requires release-specific, distance-aware verification rather than a one-time benchmark.