Applying Psychometrics to Large Language Model Simulated Populations: Recreating the HEXACO Personality Inventory Experiment with Generative Agents

๐Ÿ“… 2025-08-01
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study investigates whether large language model (LLM)-based generative agents can serve as valid substitutes for human participants in social science research, specifically evaluating their psychometric validity in representing personality traits. Method: Leveraging GPT-4, we constructed role-based generative agents and replicated the HEXACO Personality Inventory experiment using systematic prompt engineering, confirmatory and exploratory factor analysis, and cross-model comparative evaluationโ€”the first application of classical psychometric methodology to LLM-generated populations. Contribution/Results: Agent responses partially reproduce the six-factor HEXACO structure with strong internal consistency, yet exhibit significant model-specific biases. Systematic inter-model differences emerge in factor loadings and trait distributions. This work establishes a novel paradigm for empirically validating the measurability of personality in LLMs, revealing both the promise and constraints of generative agents in social science experimentation, and providing a methodological foundation and validity boundary for AI-driven behavioral simulation.

Technology Category

Cognitive Modeling & Cognitive Systems: Simulating Human BehaviorMultiagent Systems: Agent-Based Simulation and Emergent BehaviorMachine Learning: Large Multimodal Models (LMMs)

Application Category

Social Networks and Social Media: Generative AI / large language models and their impact on social systemsSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
๐Ÿ“ Abstract
Generative agents powered by Large Language Models demonstrate human-like characteristics through sophisticated natural language interactions. Their ability to assume roles and personalities based on predefined character biographies has positioned them as cost-effective substitutes for human participants in social science research. This paper explores the validity of such persona-based agents in representing human populations; we recreate the HEXACO personality inventory experiment by surveying 310 GPT-4 powered agents, conducting factor analysis on their responses, and comparing these results to the original findings presented by Ashton, Lee, & Goldberg in 2004. Our results found 1) a coherent and reliable personality structure was recoverable from the agents' responses demonstrating partial alignment to the HEXACO framework. 2) the derived personality dimensions were consistent and reliable within GPT-4, when coupled with a sufficiently curated population, and 3) cross-model analysis revealed variability in personality profiling, suggesting model-specific biases and limitations. We discuss the practical considerations and challenges encountered during the experiment. This study contributes to the ongoing discourse on the potential benefits and limitations of using generative agents in social science research and provides useful guidance on designing consistent and representative agent personas to maximise coverage and representation of human personality traits.
Problem

Research questions and friction points this paper is trying to address.

Evaluating validity of LLM agents in replicating human personality studies
Comparing GPT-4 agents' HEXACO results with original human data
Identifying model biases in generative agents' personality profiling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Using GPT-4 agents for HEXACO personality recreation
Factor analysis on agent responses for validation
Cross-model analysis reveals model-specific biases
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
S
Sarah Mercer
The Alan Turing Institute
Daniel P. Martin
Daniel P. Martin
Camulos
Data SciencePattern MatchingSide Channel Analysis
P
Phil Swatton
The Alan Turing Institute