Learning to Simulate Individuals from Macro Social Signals

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of insufficient behavioral diversity and absent reasoning supervision in individual simulation with large language models by proposing the Macro2Mind framework. This approach innovatively transforms macro-level predicted market price trajectories into supervisory signals for micro-level behavioral reasoning, jointly optimized through explicit social behavior decomposition, regret-aware curriculum sampling, and Group Relative Policy Optimization (GRPO) reinforcement learning. Experimental results demonstrate that Macro2Mind achieves state-of-the-art performance on SWM-Bench while exhibiting superior zero-shot cross-domain transferability, ultimately improving downstream simulator accuracy by 15.5 percentage points.
📝 Abstract
Large language models are increasingly used to simulate how individuals respond to new situations, yet the behavioral reasoning behind these responses is either inherited from pretraining or learned from individual-level annotations, which offer limited behavioral diversity and little supervision of the reasoning itself. We propose to learn behavioral reasoning from prediction markets, whose price trajectories record how populations respond to real-world events at scale. We introduce macro2mind, which trains a language model with GRPO using market signals. A social behavioral decomposition makes behavioral reasoning an explicit step of forecasting: the model infers representative groups of market participants, predicts how each interprets the news and updates its beliefs, reasons about their interactions, and aggregates these responses into a price. A hindsight-regret curriculum with difficulty-aware sampling focuses training on transitions where hindsight-identified groups substantially improve the forecast while prioritizing examples that remain learnable for the current policy. The learned reasoning applies to user simulation without further training. On SWM-Bench, macro2mind achieves state-of-the-art directional accuracy and correlation on Polymarket. Trained on market data, it transfers zero-shot to four user-simulation benchmarks (Humanual, OvertonBench, PRISM, and CAD) and has competitive performance among zero-shot methods. Used as a data generator, macro2mind also raises a downstream simulator's accuracy on unseen users by 15.5 points, outperforming data generated by its backbone by 13.2 points.
Problem

Research questions and friction points this paper is trying to address.

individual simulation
behavioral reasoning
large language models
behavioral diversity
Innovation

Methods, ideas, or system contributions that make the work stand out.

prediction markets
behavioral reasoning
GRPO
social behavioral decomposition
zero-shot transfer
🔎 Similar Papers
No similar papers found.
Y
Yining Zhao
University of Illinois Urbana-Champaign
B
Bushi Liu
University of Illinois Urbana-Champaign
Haofei Yu
Haofei Yu
University of Illinois Urbana-Champaign
Language AgentNatural Language Processing
Zhengyang Qi
Zhengyang Qi
Scale AI
S
Shanyong Wang
University of Illinois Urbana-Champaign
C
Chuyue Li
University of Illinois Urbana-Champaign
Y
Yuxiang Liu
Independent Researcher
Jiaxuan You
Jiaxuan You
Assistant Professor, UIUC CS
Foundation ModelsGNNLarge Language Models