Attention Sinks Are Provably Necessary in Softmax Transformers: Evidence from Trigger-Conditional Tasks

📅 2026-03-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether the "attention collapse" phenomenon in Softmax-based self-attention is inevitable and reveals its intrinsic connection to normalization constraints. By designing a trigger-condition task—where the model outputs the mean of historical representations only upon encountering a specific trigger token and zero otherwise—the authors theoretically demonstrate that Softmax normalization necessarily induces attention collapse to implement the default (zero) behavior. In contrast, non-normalized ReLU attention entirely avoids this issue. Through comparative experiments on synthetic tasks and extended scenarios using both single-head and multi-head architectures, the work empirically validates that normalization is the root cause of attention collapse, provides the first theoretical justification for its necessity under Softmax, and demonstrates the effectiveness of ReLU attention in eliminating such collapse.

Technology Category

Natural Language Processing: Safety and RobustnessComputer Vision: Adversarial Attacks & RobustnessMachine Learning: Mixture of Experts (MoE)

Application Category

User Modeling, Personalization and Recommendation: Attacks and countermeasures in recommendation systemsGraph Algorithms and Modeling for the Web: Algorithms and analysis for heterogeneous, signed, attributed, multi-relational, temporal, higher-order, and annotated Web-related graphsSecurity and Privacy: Large-scale security measurements
📝 Abstract
Transformers often display an attention sink: probability mass concentrates on a fixed, content-agnostic position. We prove that computing a simple trigger-conditional behavior necessarily induces a sink in softmax self-attention models. Our results formalize a familiar intuition: normalization over a probability simplex must force attention to collapse onto a stable anchor to realize a default state (e.g., when the model needs to ignore the input). We instantiate this with a concrete task: when a designated trigger token appears, the model must return the average of all preceding token representations, and otherwise output zero, a task which mirrors the functionality of attention heads in the wild (Barbero et al., 2025; Guo et al., 2024). We also prove that non-normalized ReLU attention can solve the same task without any sink, confirming that the normalization constraint is the fundamental driver of sink behavior. Experiments validate our predictions and demonstrate they extend beyond the theoretically analyzed setting: softmax models develop strong sinks while ReLU attention eliminates them in both single-head and multi-head variants.
Problem

Research questions and friction points this paper is trying to address.

attention sink
softmax transformers
self-attention
normalization
trigger-conditional tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

attention sink
softmax normalization
ReLU attention
trigger-conditional task
self-attention collapse
🔎 Similar Papers
No similar papers found.
Y
Yuval Ran-Milo
Tel Aviv University