π€ AI Summary
This study addresses the limitation of existing user simulation benchmarks, which predominantly rely on conversational style or self-reports and thus struggle to evaluate the behavioral fidelity of personality-driven agents. To this end, we propose APB, a benchmark that introduces a novel single-trait implicit testing mechanism to avoid interference from explicit prompting. Through synthetic persona construction, automated auditing, and expert review, APB evaluates latent personality adherence across four realistic interaction scenarios: surveys, chat, web browsing, and application use. Comprising 2,460 tasks, the benchmark reveals that even state-of-the-art models achieve a maximum full-pass rate of only 84.7%. By systematically exposing behavioral boundaries under multi-attribute and cross-modal conditions, this work establishes a new paradigm for evaluating personality consistency in large language models.
π Abstract
We introduce AgentPersonaBench (APB), a benchmark evaluating whether persona conditioning faithfully steers downstream agent behavior. While language models are increasingly deployed for persona-driven user simulation, existing benchmarks primarily evaluate conversational styling or self-reports rather than authentic behavioral fidelity. APB evaluates latent persona adherence one trait at a time, embedding each target trait within a complete synthetic profile without explicitly naming the trait or disclosing the test. Ground-truth adherence is verified strictly from observable actions across four interaction surfaces of increasing realism: survey, chat, web (interactive web environments), and app (desktop software environments). APB comprises 2,460 tasks spanning 867 traits, verified through automated audits and expert review. Our evaluation of 20 frontier model arms demonstrates that high-fidelity user simulation is already attainable: leading models achieve up to 84.7% full-pass adherence under unprompted conditions. At the same time, APB identifies clear behavioral boundaries: adherence drops across interaction modalities (only 37.9-64.3% pass all four surfaces), multi-attribute demands degrade retention, and competing model families exhibit pronounced behavioral divergence.