🤖 AI Summary
This study addresses the frequent misalignment between self-reports and behavior in large language models, investigating whether this discrepancy stems from prompt artifacts or divergent internal representations. Moving beyond black-box observation, we employ activation steering to intervene in internal representations, integrating nine vector extraction methods with psychometric behavioral tasks to systematically measure the response mechanisms of both channels during risky decision-making. Our findings reveal that self-reports and behavior are independently guided by nearly orthogonal directions, such that single interventions cannot simultaneously alter both. However, combining specific directional vectors enables synergistic or antagonistic modulation. By establishing a representation-level diagnostic methodology, this work provides an interpretable internal intervention framework for evaluating the consistency of large language models.
📝 Abstract
Self-report is an appealing low-cost probe of an LLM's dispositions, but recent work finds only selective agreement between what models report and how they behave. Prior accounts establish these patterns by prompting black-box LLMs, leaving open whether the gap is a prompting artefact or a fact about how the underlying constructs are represented internally. We investigate risk-taking, a consequential dimension of agentic decision-making, using activation steering to measure self-report and behavior under the same internal intervention. We survey nine steering-vector extraction methods spanning task-specific directives, the model's own task behavior, and dispositional descriptions at two granularities, evaluated on two behavioral tasks and two psychometric instruments across four open-weight LLMs. We find that (1) a shared internal intervention does not ensure shared responsiveness: directions built from trait descriptions move self-report but leave behavior at chance, directions built from the model's own task choices do the reverse, and only task-specific directives reach both, weakly. (2) Diagnosis dissolves that exception: removing surface confounders leaves the directives only 32% of their behavioral effect. The two channels are otherwise steered by near-orthogonal directions, each channel reachable by several independent constructions. (3) An intervention composing one behavior-moving and one self-report-moving direction moves both together; flipping one sign sets them in opposition, with reported and enacted risk pointing in opposite directions on 69-89% of flipped compositions in all models. These findings move the self-report-behavior relationship from a black-box observation to a representational one that can be inspected and controlled, motivating representational checks alongside behavioral evaluation.