assess persona reliability

Designs and implements evaluation protocols, metrics, and experiments to quantify how consistently and robustly a declared persona is expressed by generative models, including measures of stability across prompts, time, and model variants (cross-model persona testing) as well as PDG stability measurement. Builds baselines (e.g., no-persona), statistical tests, and disaggregation analyses to separate persona-driven variance from model or prompt effects and to report persona robustness and reliability.

assesspersonareliability

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.56
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study investigates whether large language model (LLM)-based generative agents can serve as valid substitutes for human participants in social science research, specifically evaluating their psychometric validity in representing personality traits. Method: Leveraging GPT-4, we constructed role-based generative agents and replicated the HEXACO Personality Inventory experiment using systematic prompt engineering, confirmatory and exploratory factor analysis, and cross-model comparative evaluation—the first application of classical psychometric methodology to LLM-generated populations. Contribution/Results: Agent responses partially reproduce the six-factor HEXACO structure with strong internal consistency, yet exhibit significant model-specific biases. Systematic inter-model differences emerge in factor loadings and trait distributions. This work establishes a novel paradigm for empirically validating the measurability of personality in LLMs, revealing both the promise and constraints of generative agents in social science experimentation, and providing a methodological foundation and validity boundary for AI-driven behavioral simulation.

Comparing GPT-4 agents' HEXACO results with original human dataEvaluating validity of LLM agents in replicating human personality studiesIdentifying model biases in generative agents' personality profiling

This study investigates the instability of persona-driven generation (PDG) in large language models when applied to multiple-choice question answering (MCQA), with a particular focus on the lack of persona consistency in non-free-text outputs. To address this, we introduce the first multidimensional stability evaluation framework encompassing performance, output, and question-level correctness, enabling a systematic analysis of how model scale, domain, prompt format, and hyperparameters influence stability. Our findings reveal that prompt format exerts a stronger effect on instability than hyperparameters such as temperature; mathematical and commonsense questions are especially prone to instability; and the best- or worst-performing personas vary significantly across configurations. These results underscore the necessity of explicitly evaluating hyperparameter-induced stability in PDG applications and demonstrate its close relationship with task accuracy.

hyperparameter sensitivityLLM instabilitymodel consistency

This study investigates whether coarse-grained marginal validation alone suffices to support large language models (LLMs) as surrogate participants in human behavioral research. We construct a fixed panel of 16 lightweight personality-conditioned GPT-4 instances and evaluate their marginal alignment with human data in repeated games, systematically analyzing the impact of prompts, phrasing, and labeling on behavioral outputs. Methodologically, we integrate a fixed-panel design, symmetric Dirichlet sensitivity analysis, finite-opportunity plug-in estimation, and exact gating tests, alongside a verifiable reproduction framework that requires no real-time API calls. Results show that three of four game units meet preregistered marginal criteria; prompt variations account for 47%–96% of behavioral variance; and minor wording adjustments increase cooperation rates from 0/40 to 37/40. While marginal matching is achievable, treatment effect estimates remain imprecise, revealing methodological limitations including household-level error, interdependence, and boundary uncertainty.

large language modelsmarginal validationpersona panels

Doubly-Robust LLM-as-a-Judge: Externally Valid Estimation with Imperfect Personas

Sep 26, 2025
LG
Luke Guerdan
🏛️ Carnegie Mellon University | Stanford University

Generative AI evaluation suffers from external validity challenges: human annotator demographics and output sample distributions in laboratory settings often deviate from real-world deployment conditions, leading to biased quality estimates. To address this, we propose a doubly robust evaluation framework that integrates large language model (LLM)-simulated, diverse annotator personas with propensity score reweighting and outcome regression modeling to yield unbiased system quality estimates. Its double robustness property—guaranteeing consistency if either the persona model or the reweighting/regression model is correctly specified—enhances reliability under distributional shift. We theoretically establish estimator consistency and empirically validate robustness across multiple bias configurations and persona fidelity levels via a persona simulation framework. Our key contribution is the first systematic integration of LLM-driven fine-grained population modeling with causal inference techniques to address generalizability limitations in GenAI evaluation.

Addresses evaluation sampling bias in GenAI system assessmentsCombines imperfect LLM persona ratings with biased human ratingsEnsures statistically valid system quality estimates for deployment

Helpful assistant or fruitful facilitator? Investigating how personas affect language model behavior

Jul 02, 2024
PH
Pedro Henrique Luz de Araujo
🏛️ University of Vienna

This study investigates how persona assignments (e.g., teacher, woman, LGBTQ+ individual) systematically influence the behavioral outputs of large language models (LLMs). Method: We conduct a large-scale empirical analysis across 12 categories comprising 162 distinct personas, evaluating responses from seven mainstream LLMs on five task domains—mathematical reasoning, historical knowledge, value alignment, and other objective/subjective benchmarks—while rigorously controlling for confounding factors via 30 prompt rewriting techniques. Contribution/Results: We provide the first evidence that persona-induced behavioral variation substantially exceeds standard prompt sensitivity; certain effects—including enhanced logical consistency under authority-related personae and richer value expression under diverse identity personae—exhibit cross-model robustness. Statistical significance is confirmed across all models and datasets, establishing persona as a potent, controllable intervention factor. These findings introduce a novel paradigm for behavior-aware model design and controllable content generation.

Compares persona-driven outputs to control and empty settingsExamines impact of 162 personas across 12 categories on LLMsInvestigates how personas affect language model behavior

Latest Papers

What's happening recently
View more

Large language models exhibit stable, personality-like behavioral patterns that influence their generalization and safety, yet effective methods for disentangling, measuring, and controlling these traits remain lacking. This work models model personality as positions in the OCEAN (Big Five) trait space and constructs, for the first time, an editable and composable personality map in weight space. By leveraging low-rank adapters (LoRA), the approach enables monotonic adjustment of specific personality dimensions and supports linear composition to form hybrid personalities. Validated through unsupervised psychometrics, LLM-based evaluators, and human judgment benchmarks across models ranging from 4B to 32B parameters, the method significantly modulates downstream safety-related behaviors—such as neuroticism affecting frustration tolerance and agreeableness influencing flattery tendencies—while preserving general capabilities, thereby establishing a principled bridge between personality measurement, model editing, and safety alignment.

behavioral patternslanguage modelsmodel safety

Traditional self-report questionnaires are susceptible to contamination from training corpora and social desirability bias, limiting their accuracy in assessing the psychological states of role-playing agents. This work proposes Generative Projective Testing (GenPT), which for the first time integrates the Thematic Apperception Test, Rorschach Inkblot Test, and Sentence Completion Test into a standardized, three-stage pipeline powered by large language models. By leveraging CharacterRAG and AnnaAgent to construct role-playing agents and employing models such as Qwen3 to generate stimuli and interpret responses, GenPT enables context-sensitive, contamination-resistant, and low-bias psychological measurement. Empirical results demonstrate that GenPT maintains a symmetric baseline under social desirability interference and captures longitudinal changes in depression indicators during counseling sessions with an order-of-magnitude greater sensitivity than conventional questionnaires, substantially enhancing both validity and stability.

projective testingpsychometric reliabilityself-report bias

This study addresses the absence of behavioral scales for precisely modulating the expression intensity of personality traits in large language models by proposing PersonaDose. This method integrates a descriptive conditioning controller (FLAS) with flow-time calibration, decoupling the controller’s learning range from query precision to enable graded control of personality traits based on target intensities without requiring paired target-intensity data. Experimental results demonstrate that PersonaDose significantly improves the expression accuracy of core personality traits on models such as Llama, achieving an average positioning error of only 4.7–6.2 points and outperforming activation addition baselines.

activation steeringcalibrationlanguage model

This study addresses the limitation of existing user simulation benchmarks, which predominantly rely on conversational style or self-reports and thus struggle to evaluate the behavioral fidelity of personality-driven agents. To this end, we propose APB, a benchmark that introduces a novel single-trait implicit testing mechanism to avoid interference from explicit prompting. Through synthetic persona construction, automated auditing, and expert review, APB evaluates latent personality adherence across four realistic interaction scenarios: surveys, chat, web browsing, and application use. Comprising 2,460 tasks, the benchmark reveals that even state-of-the-art models achieve a maximum full-pass rate of only 84.7%. By systematically exposing behavioral boundaries under multi-attribute and cross-modal conditions, this work establishes a new paradigm for evaluating personality consistency in large language models.

agent behaviorbehavioral fidelitybenchmarking

Hot Scholars

CC

Chen Chen

Assistant Professor, Computing and Information Sciences, Florida International University
Human Computer InteractionExtended Reality (XR)Healthcare
MS

Manas Satish Bedmutha

University of California San Diego
Human Computer InteractionHealthUbiquitous Computing
BH

Bo Hu

Faculty of Humanities and Arts, Macau University of Science and Technology
Human-Computer InteractionTechnology AcceptanceInteractive Media EffectsComputer-Mediated
ER

Eric Rudolph

Technische Hochschule Nürnberg
Natural Language ProcessingAI in Social WorkLarge Language ModelsVirtual Patients
PA

Phan Anh Duong

University of Cincinnati
natural language processing