Score
Designs and implements evaluation protocols, metrics, and experiments to quantify how consistently and robustly a declared persona is expressed by generative models, including measures of stability across prompts, time, and model variants (cross-model persona testing) as well as PDG stability measurement. Builds baselines (e.g., no-persona), statistical tests, and disaggregation analyses to separate persona-driven variance from model or prompt effects and to report persona robustness and reliability.
This paper systematically investigates critical challenges in large language model (LLM)-driven role-playing (RP), focusing on character authenticity and personalization. It identifies three core problems: weak personality consistency, difficulty in behavior alignment with role specifications, and insufficient user engagement. To address these, the paper proposes the first four-dimensional technical taxonomy—spanning data curation, model alignment, agent architecture, and evaluation—highlighting dynamic personality modeling and higher-order consistency as pivotal research directions. Methodologically, it integrates prompt engineering, supervised and reinforcement fine-tuning, multi-agent collaboration, and hybrid human-automated multidimensional evaluation. Key contributions include: (1) the first structured, comprehensive research landscape map for RP; (2) an open-sourced, authoritative RP literature repository on GitHub; and (3) a reproducible benchmark evaluation framework. Collectively, these advances provide both theoretical foundations and practical paradigms for immersive AI-character interaction.
This study investigates whether large language model (LLM)-based generative agents can serve as valid substitutes for human participants in social science research, specifically evaluating their psychometric validity in representing personality traits. Method: Leveraging GPT-4, we constructed role-based generative agents and replicated the HEXACO Personality Inventory experiment using systematic prompt engineering, confirmatory and exploratory factor analysis, and cross-model comparative evaluation—the first application of classical psychometric methodology to LLM-generated populations. Contribution/Results: Agent responses partially reproduce the six-factor HEXACO structure with strong internal consistency, yet exhibit significant model-specific biases. Systematic inter-model differences emerge in factor loadings and trait distributions. This work establishes a novel paradigm for empirically validating the measurability of personality in LLMs, revealing both the promise and constraints of generative agents in social science experimentation, and providing a methodological foundation and validity boundary for AI-driven behavioral simulation.
This study investigates the instability of persona-driven generation (PDG) in large language models when applied to multiple-choice question answering (MCQA), with a particular focus on the lack of persona consistency in non-free-text outputs. To address this, we introduce the first multidimensional stability evaluation framework encompassing performance, output, and question-level correctness, enabling a systematic analysis of how model scale, domain, prompt format, and hyperparameters influence stability. Our findings reveal that prompt format exerts a stronger effect on instability than hyperparameters such as temperature; mathematical and commonsense questions are especially prone to instability; and the best- or worst-performing personas vary significantly across configurations. These results underscore the necessity of explicitly evaluating hyperparameter-induced stability in PDG applications and demonstrate its close relationship with task accuracy.
This study investigates whether coarse-grained marginal validation alone suffices to support large language models (LLMs) as surrogate participants in human behavioral research. We construct a fixed panel of 16 lightweight personality-conditioned GPT-4 instances and evaluate their marginal alignment with human data in repeated games, systematically analyzing the impact of prompts, phrasing, and labeling on behavioral outputs. Methodologically, we integrate a fixed-panel design, symmetric Dirichlet sensitivity analysis, finite-opportunity plug-in estimation, and exact gating tests, alongside a verifiable reproduction framework that requires no real-time API calls. Results show that three of four game units meet preregistered marginal criteria; prompt variations account for 47%–96% of behavioral variance; and minor wording adjustments increase cooperation rates from 0/40 to 37/40. While marginal matching is achievable, treatment effect estimates remain imprecise, revealing methodological limitations including household-level error, interdependence, and boundary uncertainty.
Generative AI evaluation suffers from external validity challenges: human annotator demographics and output sample distributions in laboratory settings often deviate from real-world deployment conditions, leading to biased quality estimates. To address this, we propose a doubly robust evaluation framework that integrates large language model (LLM)-simulated, diverse annotator personas with propensity score reweighting and outcome regression modeling to yield unbiased system quality estimates. Its double robustness property—guaranteeing consistency if either the persona model or the reweighting/regression model is correctly specified—enhances reliability under distributional shift. We theoretically establish estimator consistency and empirically validate robustness across multiple bias configurations and persona fidelity levels via a persona simulation framework. Our key contribution is the first systematic integration of LLM-driven fine-grained population modeling with causal inference techniques to address generalizability limitations in GenAI evaluation.
This study investigates how persona assignments (e.g., teacher, woman, LGBTQ+ individual) systematically influence the behavioral outputs of large language models (LLMs). Method: We conduct a large-scale empirical analysis across 12 categories comprising 162 distinct personas, evaluating responses from seven mainstream LLMs on five task domains—mathematical reasoning, historical knowledge, value alignment, and other objective/subjective benchmarks—while rigorously controlling for confounding factors via 30 prompt rewriting techniques. Contribution/Results: We provide the first evidence that persona-induced behavioral variation substantially exceeds standard prompt sensitivity; certain effects—including enhanced logical consistency under authority-related personae and richer value expression under diverse identity personae—exhibit cross-model robustness. Statistical significance is confirmed across all models and datasets, establishing persona as a potent, controllable intervention factor. These findings introduce a novel paradigm for behavior-aware model design and controllable content generation.
Large language models exhibit stable, personality-like behavioral patterns that influence their generalization and safety, yet effective methods for disentangling, measuring, and controlling these traits remain lacking. This work models model personality as positions in the OCEAN (Big Five) trait space and constructs, for the first time, an editable and composable personality map in weight space. By leveraging low-rank adapters (LoRA), the approach enables monotonic adjustment of specific personality dimensions and supports linear composition to form hybrid personalities. Validated through unsupervised psychometrics, LLM-based evaluators, and human judgment benchmarks across models ranging from 4B to 32B parameters, the method significantly modulates downstream safety-related behaviors—such as neuroticism affecting frustration tolerance and agreeableness influencing flattery tendencies—while preserving general capabilities, thereby establishing a principled bridge between personality measurement, model editing, and safety alignment.
Traditional self-report questionnaires are susceptible to contamination from training corpora and social desirability bias, limiting their accuracy in assessing the psychological states of role-playing agents. This work proposes Generative Projective Testing (GenPT), which for the first time integrates the Thematic Apperception Test, Rorschach Inkblot Test, and Sentence Completion Test into a standardized, three-stage pipeline powered by large language models. By leveraging CharacterRAG and AnnaAgent to construct role-playing agents and employing models such as Qwen3 to generate stimuli and interpret responses, GenPT enables context-sensitive, contamination-resistant, and low-bias psychological measurement. Empirical results demonstrate that GenPT maintains a symmetric baseline under social desirability interference and captures longitudinal changes in depression indicators during counseling sessions with an order-of-magnitude greater sensitivity than conventional questionnaires, substantially enhancing both validity and stability.
This study addresses the absence of behavioral scales for precisely modulating the expression intensity of personality traits in large language models by proposing PersonaDose. This method integrates a descriptive conditioning controller (FLAS) with flow-time calibration, decoupling the controller’s learning range from query precision to enable graded control of personality traits based on target intensities without requiring paired target-intensity data. Experimental results demonstrate that PersonaDose significantly improves the expression accuracy of core personality traits on models such as Llama, achieving an average positioning error of only 4.7–6.2 points and outperforming activation addition baselines.
This study addresses the limitation of existing user simulation benchmarks, which predominantly rely on conversational style or self-reports and thus struggle to evaluate the behavioral fidelity of personality-driven agents. To this end, we propose APB, a benchmark that introduces a novel single-trait implicit testing mechanism to avoid interference from explicit prompting. Through synthetic persona construction, automated auditing, and expert review, APB evaluates latent personality adherence across four realistic interaction scenarios: surveys, chat, web browsing, and application use. Comprising 2,460 tasks, the benchmark reveals that even state-of-the-art models achieve a maximum full-pass rate of only 84.7%. By systematically exposing behavioral boundaries under multi-attribute and cross-modal conditions, this work establishes a new paradigm for evaluating personality consistency in large language models.
为了解决工具增强LLM代理评估中用户输入缺乏多样性的问题,本文提出了一种包含23个维度的三层人格向量模型,以生成更真实多样的用户模拟。