Benchmarking Psychological Dynamics in Generative Agents

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of benchmarks and internal validity in using large language models (LLMs) to simulate human psychological dynamics by constructing the first temporally evolving dynamic personality benchmark that jointly ensures internal and external validity. Employing psychometric paradigms, meta-analytic effect sizes, and variance decomposition techniques, this work systematically evaluates the psychological realism of 36 LLMs regarding trajectory distributions and trait relationships. Results indicate that while most models successfully recover the directionality of trait relationships, they struggle to replicate within-person correlations. Furthermore, model scale is found to be independent of internal validity, with Gemma-3-27B achieving the optimal performance trade-off.
📝 Abstract
Large language models (LLMs) are increasingly deployed to simulate human behavior, acting as computational replicas of human subjects. Yet the lived psychological experience of humans is difficult to benchmark, particularly as it unfolds over time. We introduce a psychometric benchmark for computational replicas: personas that carry a fixed identity through an evolving sequence of events. Built entirely from published norms and meta-analytic effects, the benchmark scores two dimensions of psychological realism. The first, internal validity, quantifies whether generated trajectories reproduce the internal structure of repeated human measurement: distributions, the between- versus within-person variance partition, temporal dependence, and range. The second, external validity, quantifies whether replicas recover established trait, state, and indicator relations. Across 36 open-weight and proprietary LLMs from nine developers (1B-671B parameters), most recover the direction of established relations (84.3% mean agreement) and the variance partition (26 of 34), yet the within-person correlation and distributional structure elude recovery. Internal validity is independent of scale and capability: a mid-size open LLM (Gemma-3-27B) strikes the best trade-off between the two dimensions. The benchmark is a precondition for using computational replicas in causal inference across domains (e.g., marketing, healthcare), and identifies within-person grounding as the central challenge ahead.
Problem

Research questions and friction points this paper is trying to address.

Generative Agents
Psychological Dynamics
Large Language Models
Benchmarking
Computational Replicas
Innovation

Methods, ideas, or system contributions that make the work stand out.

Psychometric Benchmark
Generative Agents
Internal Validity
Within-person Variance
Computational Replicas
💼 Related Jobs
No related jobs found.