Score
Designs and conducts analyses of individual attention heads within transformer models, characterizing their specialization, representational content, and causal contribution to overall model behavior using probing, interventions, and ablation tests. Tasks include detecting compromised or safety-aligned heads, attributing effects to specific tokens or mechanisms, and quantifying each head’s functional role.
This work addresses the limited trustworthiness of Transformer models in high-stakes applications, which stems from insufficient understanding of their internal decision-making mechanisms. To bridge this gap, we propose a mechanistic interpretability approach based on targeted interventions on attention heads, integrating causal analysis with neural circuit probing to systematically uncover the model’s decision processes and underlying cognitive mechanisms. Our method substantially enhances the interpretability of Transformer internals and offers an innovative pathway toward the design and control of highly reliable AI systems, while also enabling the discovery of novel scientific insights encoded within these models.
Existing interpretability methods often erroneously attribute specific computational roles to attention heads without verifying their generalization across diverse prompts. This work proposes the KID role classification framework and a three-stage analysis pipeline that combines activation patching with same-answer control conditions to expose widespread pseudo-semantic specificity in conventional attribution approaches. By integrating capability-selective screening (CSS), singular value decomposition (SVD), and activation transduction under matched controls, we systematically evaluate attention head functionality across multiple 7–8B instruction-tuned models. Our findings demonstrate that the majority of attention heads previously identified by standard methods as performing specific roles fail to consistently transfer their purported computational functions across different prompts, thereby challenging the dominant attribution paradigm in mechanistic interpretability.
提出一种影响评分方法,量化注意力头在基于Transformer的模型中对分类决策的贡献,用于提升模型在提示注入检测中的解释性。
This work investigates how attention heads in small Transformers collaborate during counting tasks: do they operate via pseudo-ensemble redundancy or functionally differentiated division of labor? Using mechanistic interpretability analysis, attention pattern visualization, and head-level functional attribution, we find— for the first time—that heads exhibit high semantic redundancy, jointly executing identical subtasks (i.e., pseudo-ensemble behavior); yet syntactic correctness critically depends on non-uniform weighted aggregation of head outputs to satisfy grammatical constraints. This reveals a fundamental decoupling between semantic redundancy and syntactic sensitivity within the attention mechanism. Our results challenge assumptions about functional specialization in small Transformers and establish a new paradigm for analyzing their internal structure and interpretability.
This work investigates the learnable solution space of small Transformers when solving histogram tasks—i.e., counting token frequencies in sequences—a deceptively simple yet revealing probe of model internals. Methodologically, the authors integrate theoretical modeling (via linear-algebraic characterization), empirical training, and mechanistic reverse-engineering (including attention visualization and gradient probing). They formally distinguish, for the first time, two distinct counting strategies: relation-based and inventory-based. Results show fine-grained coordination between attention and feed-forward networks; minor architectural changes (e.g., replacing softmax) induce abrupt strategy shifts. Quantitative analysis uncovers nonlinear performance boundaries governed by vocabulary size, embedding dimension, FFN capacity, and attention design. Crucially, both strategies emerge empirically during training, with their prevalence determined jointly by hyperparameters and implicit couplings among model components.
Bias in Transformer language models is challenging to precisely locate and mitigate. This work proposes ROBIN, a method that performs sensitivity analysis during inference using fairness probes to identify and rank attention heads associated with bias. Rather than naively zeroing out entire heads, ROBIN precisely removes bias directions within the output subspaces of these identified heads. This approach enables white-box fairness debugging at the granularity of individual attention heads. Evaluated across four mainstream Transformer models, ROBIN significantly reduces the WinoBias gap while outperforming naive head-zeroing strategies and better preserving overall language modeling performance.
研究通过RFIS和RPD指标分析RoPE-based Transformers的功能组织,提出Head-wise Hybrid Architecture (HwH),用NoPE FA进行全局检索、LA进行局部位置建模,以改善零样本长上下文外推性能。
本文证明了多头自注意力机制可视为参数识别策略,探讨了更多头数如何提高模型的识别度,并通过数学和实验验证了这一理论。
This study investigates how language models internally represent network infrastructure information, such as hostname-IP pairs. Through causal ablation and intervention experiments, the authors localize and validate specific attention heads and neurons responsible for this task, evaluating their generalizability via correlation-based ranking and cross-dataset transfer tests. The findings demonstrate that attention-head-level causal mechanisms exhibit cross-architecture universality, with full-head detectors achieving 99.5%–100% accuracy across multiple models. Conversely, neuron-level responsibility distributions are shown to be model-specific, limiting the transferability of single-neuron approaches and necessitating per-model validation.
This study investigates how large language models detect internal activation perturbations and localize the positions of such changes. Through concept vector injection, attention head intervention, QK/OV computation analysis, and comparative experiments across multiple model families, it systematically dissects the underlying mechanisms of attention heads. The work makes two primary contributions: first, it identifies for the first time distinct gating heads responsible for detecting changes and routing heads responsible for selecting positions, revealing their mutually inhibitory interaction; second, it elucidates the relationship between localization precision and attention response magnitude, establishing the micro-level mechanisms that support introspective detection within these models.