Identifying Introspection From the Inside

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of distinguishing genuine introspection from confabulated self-reports in large language models by examining their internal mechanisms. It proposes a novel fidelity detection paradigm grounded in network architecture rather than solely behavioral observation. By employing low-rank adaptation to train models for decision simulation, the authors conduct mechanistic analyses using weight ablation, layer-freezing experiments, and attribution patching. The findings reveal that faithful self-reporting is characterized by the migration of preference representations toward shallower layers and high cross-task attribution similarity. This work uncovers the mechanistic signatures of model introspection, demonstrating that faithful self-reports are accompanied by specific structural changes. Ultimately, it enables the identification of divergent computational patterns without requiring semantic interpretation of model content.
📝 Abstract
Large language models make claims about themselves that are both consequential and increasingly difficult to verify from behavior alone. How can we distinguish plausible confabulations from genuine introspection? In this paper, we identify mechanistic signatures of faithful self-report in a controlled setting. Using low-rank adapters, we train models to make decisions on behalf of fictitious characters, according to latent linear preference functions. We find sustained fine-tuning on an implicit decision task can lead to the emergence of accurate self-reporting of models' learned preferences, even without explicit self-report supervision. We ask two research questions about this emergent phenomenon. First: is the emergence of accurate self-reporting accompanied by a measurable structural change in the model? Weight ablations and frozen-layer experiments together indicate that preference representations shift to earlier layers over training, consistent with the hypothesis that faithful self-report requires preferences to be located where pre-existing verbalization mechanisms can access them. Second: can these structural differences distinguish faithful models from unfaithful ones? Using attribution patching, we find that faithful models exhibit significantly higher attribution similarity between the decision-making and self-report tasks -- a mechanistic signature of faithful self-report that does not require us to understand the content of the report itself. Previous work on self-report has observed behaviorally that models can be faithful or unfaithful; our work proposes that, at least in our restricted setting, it is possible to distinguish between the two patterns of computation by examining the structure of the networks themselves.
Problem

Research questions and friction points this paper is trying to address.

large language models
introspection
faithful self-report
confabulation
mechanistic interpretability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Introspection
Low-rank adapters
Attribution patching
Mechanistic interpretability
Self-report
🔎 Similar Papers
No similar papers found.