Exact Attention Sensitivity and the Geometry of Transformer Stability

📅 2026-02-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the long-standing lack of theoretical understanding behind the training instability of Transformers, particularly regarding the interplay among normalization schemes, learning rate warmup, and depth scaling. Building from first principles, the authors develop a stability theory grounded in attention sensitivity and the geometric structure of gradient flow. They derive the exact operator norm of the softmax Jacobian for the first time, introduce a sequence-length-independent balance quality factor θ(p) to quantify attention sensitivity, and establish corresponding Lipschitz bounds. Through operator norm analysis, block-wise ∞/RMS geometric modeling, and gradient path characterization, they demonstrate that architectural design—not learned attention patterns—governs stability. The theory successfully explains the efficacy of pre-LayerNorm, DeepNorm’s N⁻¹/⁴ scaling law, and the necessity of learning rate warmup. Empirical validation on a 774M-parameter model confirms θ(p)≈1 throughout training, underscoring that stability is inherently architectural.

Technology Category

Machine Learning: Deep Neural Architectures and Foundation ModelsNatural Language Processing: Learning & Optimization for NLPSearch and Optimization: Learning to Search

Application Category

Security and Privacy: Large-scale security measurementsSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsGraph Algorithms and Modeling for the Web: Algorithms and analysis for heterogeneous, signed, attributed, multi-relational, temporal, higher-order, and annotated Web-related graphs
📝 Abstract
Despite powering modern AI, transformers remain mysteriously brittle to train. We develop a stability theory that explains why pre-LayerNorm works, why DeepNorm uses $N^{-1/4}$ scaling, and why warmup is necessary, all from first principles. Our framework has two pillars: (1) We derive the \emph{exact} operator norm of the softmax Jacobian, $\|J_{softmax}(u/τ)\|_{\infty\to 1} = θ(p)/τ$, where the balanced-mass factor $θ(p)\in[0,1]$ quantifies attention sensitivity. (2) We introduce a block-$\infty$/RMS geometry aligned with tokenwise computation, yielding Lipschitz bounds independent of sequence length. Using this framework, we prove that pre-LN preserves identity gradient paths while post-LN compounds LayerNorm Jacobians exponentially with depth, and we show that DeepNorm's $N^{-1/4}$ emerges from the quartic structure of attention's four projection matrices. We validate our theory on 774M-parameter models and find that, contrary to the intuition that attention sharpens during training to reduce sensitivity, $θ(p) \approx 1$ persists throughout. Transformer stability arises entirely from architectural gradient flow, not from attention dynamics. This finding changes how we reason about training: the architecture itself must handle sensitivity, not learned attention patterns.
Problem

Research questions and friction points this paper is trying to address.

Transformer stability
attention sensitivity
training brittleness
softmax Jacobian
gradient flow
Innovation

Methods, ideas, or system contributions that make the work stand out.

attention sensitivity
softmax Jacobian
transformer stability
pre-LayerNorm
Lipschitz bound
🔎 Similar Papers