Score
Designs and implements attention mechanisms that compute and apply element- or feature-wise gating weights to attention scores, attention outputs, or cross-layer signals (including per-dimension, channel-/spectral, region-gated, latent-query, self- and cross-attention variants) to selectively modulate information flow. These modules compress or route encoder/decoder or inter-layer representations, suppress irrelevant or noisy features, provide fine-grained control of instance contributions, and enable more efficient or lightweight attention computation.
Deep neural networks lack biologically plausible selective attention mechanisms, limiting both efficiency and accuracy in image recognition. To address this, we propose a spatial attention module grounded in information bottleneck theory. Our method explicitly optimizes mutual information: it minimizes the mutual information between the attention representation and the input to suppress redundancy, while maximizing the mutual information between the attention representation and task labels to enhance discriminability. Crucially, we introduce learnable anchors to quantize continuous attention scores—a novel design that strengthens information constraints and improves interpretability of attention maps. By integrating variational attention modeling with deep network embedding, our approach achieves significant performance gains across image classification, fine-grained recognition, and cross-domain classification tasks. The resulting attention maps exhibit high discriminability, strong background suppression, and enhanced interpretability.
研究探讨了在稀疏注意力机制中,学习到的门控与随机门控的效果差异,通过控制实验揭示模型表示与施加掩码的共适应导致学习门控优势有限。
This work addresses the challenge of enhancing neural networks’ ability to focus on salient information in long-sequence and multimodal tasks. By establishing a unified theoretical framework for attention mechanisms, the study systematically analyzes their mathematical foundations, computational properties, and cross-task generalizability. The framework is instantiated across diverse architectures—including autoregressive Transformers, bidirectional encoders, Vision Transformers, and cross-modal attention models—demonstrating consistent performance gains. The research further uncovers an intrinsic relationship between attention structure and model interpretability, validates empirical scaling laws governing training dynamics and performance, and achieves state-of-the-art results on multiple benchmark datasets. Attention visualization techniques are employed to enhance model transparency, offering insights into the decision-making process of these architectures.
This work addresses the poor conditioning of the Jacobian matrix in Transformer attention mechanisms, which often leads to training instability and performance degradation. For the first time, it explicitly establishes a theoretical link between the condition number of the attention Jacobian and the spectral properties of the query, key, and value projection matrices. Building on this insight, the paper proposes a general, plug-and-play spectral regularization strategy that improves the Jacobian’s condition number by optimizing the singular value distribution of these projection matrices. Notably, the method requires no architectural modifications and consistently enhances performance across diverse Transformer variants and tasks, demonstrating both its effectiveness and broad applicability.
This work addresses the lack of a systematic understanding of how attention mechanisms in Vision Transformers jointly process positional and content information. The authors propose a Bilinear Factorization Decomposition (BFD) framework that, for the first time, achieves statistical disentanglement of positional and content factors through ANOVA decomposition, combined with singular value decomposition (SVD) of the QK^T matrix to uncover dominant interaction modes within attention. Their analysis reveals that attention energy is primarily driven by content-content interactions; DINOv2 exhibits stronger content-position coupling and a richer distribution of interaction modes; and intermediate layers enhance shape perception by jointly preserving positional structure and amplifying semantic signals.
This work addresses the attention sink problem in attention mechanisms and the inherent trade-off between representational capacity and computational efficiency by proposing a Hybrid Gated Attention (HyGA) framework. HyGA introduces, for the first time, a multi-view, multi-stage hybrid gating mechanism that integrates learnable attention sinks, low-rank matrix decomposition, and element-wise and head-wise collaborative modulation to enable fine-grained control over information flow. The proposed method consistently outperforms existing gated attention approaches across diverse backbone architectures and benchmark tasks, achieving state-of-the-art performance under varying computational budgets. Furthermore, HyGA effectively reduces training loss while simultaneously enhancing model stability and expressive power.
This work addresses the entanglement of routing and filtering functions in conventional attention mechanisms, which leads to structural opacity and optimization challenges. The authors propose S-D Attention, which explicitly decouples the interaction matrix into a low-rank routing component and a symmetric filtering component, thereby clearly distinguishing these two mechanisms for the first time. They further uncover that routing self-organizes into a spectral cascade phenomenon in deep networks. Leveraging this insight, they achieve stable training without layer normalization and validate their approach using linear attention variants (e.g., ELU+1) and effective rank analysis. Experiments show that linearizing the first seven layers of a 125M-parameter model incurs less than 5% perplexity degradation, while cascade-informed architectures reduce attention parameters by 47%–65% with only a 3.9%–8.4% increase in perplexity.
This work addresses the commonly oversimplified notion of “attention sinks” in attention mechanisms, revealing that they actually encompass two distinct computational phenomena: adaptive no-ops (nop) and broadcast operations. The study formally distinguishes these mechanisms for the first time and introduces verifiable diagnostic criteria based on value vector norms and output rank. Through synthetic tasks and empirical analysis of pretrained vision Transformers, the authors demonstrate their coexistence and layer-wise distribution patterns. Furthermore, the research shows that existing intervention strategies—such as gating and register tokens—implicitly favor one mechanism over the other, and that combining both approaches yields complementary benefits, leading to improved model stability and performance.
This work addresses the issue of diffuse attention distributions in self-attention mechanisms, which often undermine model interpretability. By analyzing the parametric structures of the query–key and output–value circuits and their differing learning rates, the study reveals that the relative learning rate critically governs attention sharpness. Through gradient flow analysis and closed-form dynamical derivations, the authors theoretically demonstrate this relationship and empirically validate it using a single-layer self-attention model. The experiments show that assigning a higher learning rate to the query–key circuit naturally induces sharper, more concentrated attention patterns. This approach significantly improves interpretability metrics while maintaining comparable predictive performance, offering a novel pathway to enhance attention focus without requiring additional regularization.
This work addresses the limitations of existing layer attention mechanisms—namely high computational complexity, static information updating, and inadequate modeling of long-range dependencies—by proposing Key-Correlated Layer Attention (KCLA). KCLA leverages the high cosine similarity among inter-layer Key representations to establish a dynamic cross-layer interaction mechanism with linear time complexity and constant space complexity. By integrating key-correlation-based linear attention with adaptive information fusion, KCLA maintains strong long-range dependency modeling capabilities while achieving computational efficiency independent of network depth. Experimental results demonstrate consistent and significant performance improvements across diverse tasks, including image classification, object detection, and medical image segmentation.