On the Behavioral Traits of LLM Agents

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of self-report bias and high manual annotation costs in AI personality assessment by proposing the A-B-D framework, which pioneers a bottom-up quantitative approach to infer agent traits from authentic interaction behaviors. Through feature engineering and factor analysis on 340,000 trajectory records, functional and linguistic features were extracted, revealing six stable personality factors. The research uncovers a pronounced "attitude-behavior gap" in AI agents: it successfully identifies trait-level differences among models such as Kimi and demonstrates that behavioral factors exhibit negligible correlation with self-reported Big Five personality scores. These findings establish a novel paradigm for the objective evaluation of AI personality.
📝 Abstract
Users increasingly describe different AI agents as distinct colleagues to work with. AI personality research aims to quantify such impressions by attributing human-like"traits"to agents. However, existing measures fall short: models'self-reports (S-data) diverge from their actual behavior, while informant ratings from LLM judges (I-data) are costly to scale and cover few everyday scenarios. In this paper, we propose A-B-D to infer traits bottom-up from behavioral data (B-data), namely how agents act on their environment and communicate with users, as recorded in existing trajectories. From 345,667 real-world trajectories spanning 80 models, 12 tasks, and 50 harnesses, we extract 318 candidate features that capture both the actions an agent takes at each step (functional) and the language accompanying them (linguistic). We retain only features that show instance-level stability, cross-task consistency, and model discriminability. Factor analysis of the remaining 79 features uncovers six stable, model-attributable factors, two functional and four linguistic. For example, Kimi-K3 exhibits the most planfulness, whereas GPT-5.5 and GPT-5.6 are the least energetic. Moreover, we quantify the"knowledge-action gap"in the wild: these factors correlate only weakly with self-reported Big Five scores, even for conceptually matched pairs such as extroversion and energetic (r = 0.07, p = 0.58). Our work offers a new lens for understanding AI personality, with implications for users, developers, and researchers from both computer science and social science.
Problem

Research questions and friction points this paper is trying to address.

LLM agents
AI personality
behavioral traits
self-report divergence
knowledge-action gap
Innovation

Methods, ideas, or system contributions that make the work stand out.

Behavioral Traits
Bottom-up Inference
Factor Analysis
Knowledge-Action Gap
LLM Agents
🔎 Similar Papers