🤖 AI Summary
This study addresses the reliance on costly target-domain data and associated privacy risks when using large language models (LLMs) to simulate surveys. We propose a method that constructs demographic group profiles using anonymized public behavioral data. Through systematic analysis of how cross-domain data sources affect simulation alignment, we reveal that mismatched population distributions degrade generalization performance. Our findings demonstrate that precise profile assignment combined with rich historical behavioral data substantially enhances simulation quality. Optimized profiles achieve more stable simulation alignment on unseen questions, effectively reducing traditional survey costs. This work establishes an efficient and secure paradigm for LLM-driven computational social science.
📝 Abstract
However, many existing steering approaches rely on target-domain human data for fine-tuning or prompting that is costly to collect and raises privacy concerns. In this paper, we study demographic group-level survey simulation, where personas induced from heterogeneous, anonymized public behavioral data condition agents that simulate responses of individuals from specific demographic groups. We examine whether representative personas can be induced from diverse sources and analyze how the domain, scale, and granularity of the source data affect survey simulation alignment. We find that personas induced from out-of-domain sources rarely outperform simulations conditioned only on basic demographic information, largely due to population mismatch. However, when personas are accurately assigned to the target demographic groups, alignment improves substantially. Finally, personas induced from target-domain survey data generalize better as more survey question history becomes available, suggesting that richer behavioral evidence enables more stable persona trait inference that transfers to better unseen questions simulation alignment.