Score
Designs and applies analyses and tests to identify attention heads in transformer models that retrieve or route information into the model’s output distribution, including heads that copy or surface non-literal sources. This includes computing each head’s contribution to answer logits (e.g., via projection or logit-contribution scoring), ranking heads by that contribution, running needle-in-a-haystack contrasts between source positions, using write-aware detection, and selecting heads for targeted ablation.
提出一种影响评分方法,量化注意力头在基于Transformer的模型中对分类决策的贡献,用于提升模型在提示注入检测中的解释性。
Existing methods struggle to identify attention heads in large language models that perform non-literal retrieval—i.e., synthesizing answers based on semantic composition—because they focus solely on read locations while neglecting the write role of output-value (OV) circuits. This work proposes a write-aware Logit-Contribution Scoring (LOCOS) method that precisely localizes critical attention heads in a single forward pass by quantifying each head’s contribution to the logit along the direction of the target token embedding. Experiments on models such as Qwen3, Gemma-3, and OLMo-3.1 demonstrate that ablating LOCOS-identified heads causes a drastic drop in ROUGE-L scores (e.g., from 0.401 to 0.000 in Qwen3-8B), substantially outperforming baseline approaches, while preserving the model’s capacity for parameter memorization and arithmetic reasoning.
This work investigates how attention heads in small Transformers collaborate during counting tasks: do they operate via pseudo-ensemble redundancy or functionally differentiated division of labor? Using mechanistic interpretability analysis, attention pattern visualization, and head-level functional attribution, we find— for the first time—that heads exhibit high semantic redundancy, jointly executing identical subtasks (i.e., pseudo-ensemble behavior); yet syntactic correctness critically depends on non-uniform weighted aggregation of head outputs to satisfy grammatical constraints. This reveals a fundamental decoupling between semantic redundancy and syntactic sensitivity within the attention mechanism. Our results challenge assumptions about functional specialization in small Transformers and establish a new paradigm for analyzing their internal structure and interpretability.
This work investigates inter-layer information propagation in Transformer language models, focusing on how features are encoded, routed, and form cross-layer communication channels within low-rank subspaces. We identify and empirically validate the existence of “position-indexed 3D subspaces,” revealing that “contextual item crowding” is the root cause of failure in sequential sensitivity across multiple items. Methodologically, we integrate singular value decomposition (SVD), residual stream subspace analysis, low-rank feature tracking, and intervention experiments on a synthetic task (Laundry List). Crucially, we achieve the first interpretable weight editing and representation intervention grounded in this subspace structure: on the Laundry List task, accuracy improves by over 20%; we successfully predict cross-layer attention interactions; and we provide faithful, mechanistic attributions for model failures.
This work addresses the challenge of interpreting attention head functionality in large language models (LLMs), where conventional methods rely on costly inference or training. We propose MAPS, a parameter-only framework that infers head-level functional semantics without forward passes—leveraging geometric analysis in parameter space, decomposition of attention weights, functional template matching, and causal intervention validation. MAPS is the first method to enable purely parameter-driven functional mapping of attention heads, supporting quantitative measurement of operational strength and identification of individual head functionality. It uncovers previously overlooked functional patterns, revealing both functional universality across models and architecture-specific biases. Evaluated on six mainstream LLMs, MAPS achieves high correlation (>0.85) for 20 distinct operations; human evaluation confirms >90% of its functional descriptions as semantically reasonable. Furthermore, it systematically characterizes cross-model functional distribution patterns.
This work addresses the phenomenon of "attention concentration" in Transformer models, wherein a disproportionate amount of attention is allocated to uninformative or specific tokens, thereby undermining model interpretability, destabilizing training and inference, and exacerbating hallucination issues. The paper presents the first comprehensive survey of this phenomenon, introducing a three-dimensional classification framework—comprising foundational utilization, mechanistic explanation, and mitigation strategies—to systematically organize the evolving research landscape. By synthesizing recent findings on anomalous attention behaviors, the study constructs a structured knowledge base that clarifies core concepts and key challenges. It offers both theoretical insights and practical pathways for understanding and mitigating attention concentration, and further supports community advancement by releasing a curated list of relevant publications.
本文证明了多头自注意力机制可视为参数识别策略,探讨了更多头数如何提高模型的识别度,并通过数学和实验验证了这一理论。
本文提出了一种重要性评分方法来分析多头Transformer模型在表格数据学习中的作用,通过实验验证了该方法能有效提高模型效率和减少冗余。
Existing interpretability methods often erroneously attribute specific computational roles to attention heads without verifying their generalization across diverse prompts. This work proposes the KID role classification framework and a three-stage analysis pipeline that combines activation patching with same-answer control conditions to expose widespread pseudo-semantic specificity in conventional attribution approaches. By integrating capability-selective screening (CSS), singular value decomposition (SVD), and activation transduction under matched controls, we systematically evaluate attention head functionality across multiple 7–8B instruction-tuned models. Our findings demonstrate that the majority of attention heads previously identified by standard methods as performing specific roles fail to consistently transfer their purported computational functions across different prompts, thereby challenging the dominant attribution paradigm in mechanistic interpretability.
This work addresses the computational redundancy in standard multi-head attention mechanisms, which uniformly activate all attention heads regardless of task requirements or input complexity. To overcome this limitation, the authors propose BudgetFormer, the first framework that enables input-dependent dynamic head budget allocation and selection. Specifically, it employs adaptive multi-head attention to dynamically determine the optimal number of heads for each input and selects the most informative ones. An exploration-exploitation balanced training strategy is further introduced to optimize resource allocation. Extensive experiments on multiple text classification benchmarks demonstrate that BudgetFormer significantly reduces FLOPs and memory consumption while achieving performance comparable to or better than full-head attention models.