🤖 AI Summary
This study addresses the lack of systematic validation regarding the accuracy of large language models (LLMs) in simulating human belief updating. For the first time, it conducts a one-to-one comparison between six LLMs and 391 human participants on belief dynamics after reading Reddit comments, employing demographic and personality-based user profiling in prompt engineering for multi-model evaluation. Results reveal that only Qwen3-32B and GPT-5-Mini align with human posterior belief distributions when the true initial stance is known. Critically, all models fail to reliably generate plausible initial beliefs or produce credible belief updates, exhibiting three consistent biases: a tendency toward neutrality, frequent minor shifts in belief, and an inability to correctly rank comment persuasiveness. These findings highlight fundamental limitations of current LLMs as proxies for human agents in social science research.
📝 Abstract
LLMs are increasingly deployed as proxies for human study participants in social science experiments, yet the fidelity of this practice has rarely been tested directly. We test whether six LLMs can simulate individual human belief updates, comparing LLM outputs 1-to-1 against ground truth data from 391 UK participants on Prolific, who updated their stances on three discussion topics after reading Reddit comments. Each participant was simulated by an LLM conditioned on a persona derived from their demographic and personality trait data. We find that some LLMs (Qwen3-32B and GPT-5-Mini) can match the human post-stance distribution, but only when given participants' actual initial stances. All six models fail to simulate initial stances themselves and to produce faithful belief updates from self-generated stances. Three systematic biases emerge across all models: overrepresentation of neutral positions, more frequent but smaller belief shifts than humans, and a failure to rank comments by convincingness. Demographic and personality trait personas had no consistent effect on fidelity. LLM simulations of human belief dynamics are only reliable when grounded in realistic starting conditions, that current multi-round social media simulations rarely provide.