A mechanistic study of language model introspection

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates how large language models detect internal activation perturbations and localize the positions of such changes. Through concept vector injection, attention head intervention, QK/OV computation analysis, and comparative experiments across multiple model families, it systematically dissects the underlying mechanisms of attention heads. The work makes two primary contributions: first, it identifies for the first time distinct gating heads responsible for detecting changes and routing heads responsible for selecting positions, revealing their mutually inhibitory interaction; second, it elucidates the relationship between localization precision and attention response magnitude, establishing the micro-level mechanisms that support introspective detection within these models.
📝 Abstract
Large language models (LLMs) can sometimes report perturbations to their internal activations---even when the input provides no evidence that an intervention occurred. How do models detect and localize such internal changes? We study this question using a controlled task that keeps the input text fixed. We either inject a concept vector into the hidden state at one of ten token positions or apply no intervention. The model is asked to identify the perturbed position or report that no intervention occurred. Across three model families, we identify two small groups of attention heads with distinct roles in introspective reporting. Middle-layer gate heads influence whether the model reports a change, while router heads in a later layer help select the position to report. Interventions on gate heads can suppress position reports even when router heads supply location information. We further examine why reporting accuracy varies across concepts. Concept vectors that are localized more accurately produce stronger attention-score and output responses in gate heads, which is associated with better alignment of the induced key and value changes in their QK and OV computations. Together, these findings identify attention-head mechanisms supporting introspective detection and localization.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Introspection
Internal Perturbations
Attention Heads
Mechanistic Interpretability
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM introspection
attention head mechanisms
concept vectors
mechanistic interpretability
internal perturbation detection
🔎 Similar Papers
No similar papers found.