From Representations to Behaviors: Exploring the Person-Situation-Behavior Triad in LLMs

πŸ“… 2026-07-29
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the lack of empirical evidence regarding the existence of personality-related internal representations in large language models (LLMs) and their behavioral influence mechanisms. Building upon Funder’s tripartite framework of personality, the work conceptualizes Person, Situation, and Behavior as internal model representations, contextual inputs, and behavioral outputs, respectively, thereby establishing the first controllable and verifiable personality-like representation framework within LLMs. Through sparse autoencoder (SAE) decomposition, contrastive behavioral pair analysis, feature-level interventions, and token-level activation probing, the study successfully identifies sparse internal features associated with personality traits. This enables bidirectional behavioral modulation across contexts and replicates the trade-off patterns observed in human personality research within social intelligence tasks, revealing causal links among internal states, situational contexts, and behavioral outcomes.
πŸ“ Abstract
Human personality theories characterize traits not as isolated attributes captured by a single score, but as stable individual tendencies expressed through the interplay among persons, situations, and behaviors. Existing studies of personality-related behavior in LLMs have primarily focused on outputs elicited under personality conditioning, characterizing observable trait-related expressions while lacking mechanistic evidence for the existence of internal personality-related representations, their cross-situational expression, and how these representations shape specific behaviors. Building on Funder's personality triad framework, we adapt its three components for LLM analysis: Person as personality-related internal representations, Situation as contexts that afford trait-relevant responses, and Behavior as response patterns on broader social tasks. We introduce a framework for discovering, controlling, and validating trait-like representations in LLMs. First, using contrastive behavior pairs grounded in shared situations, we identify sparse internal features associated with opposing poles of personality traits through SAE decomposition. We validate their trait relevance through effects on behavior to situation, token-level activation patterns, and robustness to paraphrasing. Second, feature-level interventions induce bidirectional trait-related shifts across a separate, diverse set of situations while preserving response validity, demonstrating consistent expression across contexts. Third, applying the same interventions to social intelligence tasks reveals behavioral changes with benefit-tradeoff patterns consistent with findings from human personality research, providing behavioral-level validation beyond personality scores. Our findings provide evidence that LLMs contain controllable trait-like representations linking internal states, situational expression, and behavioral outcomes.
Problem

Research questions and friction points this paper is trying to address.

personality representations
person-situation-behavior triad
large language models
trait expression
behavioral consistency
Innovation

Methods, ideas, or system contributions that make the work stand out.

personality representations
sparse autoencoders
feature-level interventions
person-situation-behavior triad
large language models