🤖 AI Summary
This study investigates whether the "attention collapse" phenomenon in Softmax-based self-attention is inevitable and reveals its intrinsic connection to normalization constraints. By designing a trigger-condition task—where the model outputs the mean of historical representations only upon encountering a specific trigger token and zero otherwise—the authors theoretically demonstrate that Softmax normalization necessarily induces attention collapse to implement the default (zero) behavior. In contrast, non-normalized ReLU attention entirely avoids this issue. Through comparative experiments on synthetic tasks and extended scenarios using both single-head and multi-head architectures, the work empirically validates that normalization is the root cause of attention collapse, provides the first theoretical justification for its necessity under Softmax, and demonstrates the effectiveness of ReLU attention in eliminating such collapse.
📝 Abstract
Transformers often display an attention sink: probability mass concentrates on a fixed, content-agnostic position. We prove that computing a simple trigger-conditional behavior necessarily induces a sink in softmax self-attention models. Our results formalize a familiar intuition: normalization over a probability simplex must force attention to collapse onto a stable anchor to realize a default state (e.g., when the model needs to ignore the input). We instantiate this with a concrete task: when a designated trigger token appears, the model must return the average of all preceding token representations, and otherwise output zero, a task which mirrors the functionality of attention heads in the wild (Barbero et al., 2025; Guo et al., 2024). We also prove that non-normalized ReLU attention can solve the same task without any sink, confirming that the normalization constraint is the fundamental driver of sink behavior. Experiments validate our predictions and demonstrate they extend beyond the theoretically analyzed setting: softmax models develop strong sinks while ReLU attention eliminates them in both single-head and multi-head variants.