🤖 AI Summary
This work addresses the long-standing lack of theoretical understanding behind the training instability of Transformers, particularly regarding the interplay among normalization schemes, learning rate warmup, and depth scaling. Building from first principles, the authors develop a stability theory grounded in attention sensitivity and the geometric structure of gradient flow. They derive the exact operator norm of the softmax Jacobian for the first time, introduce a sequence-length-independent balance quality factor θ(p) to quantify attention sensitivity, and establish corresponding Lipschitz bounds. Through operator norm analysis, block-wise ∞/RMS geometric modeling, and gradient path characterization, they demonstrate that architectural design—not learned attention patterns—governs stability. The theory successfully explains the efficacy of pre-LayerNorm, DeepNorm’s N⁻¹/⁴ scaling law, and the necessity of learning rate warmup. Empirical validation on a 774M-parameter model confirms θ(p)≈1 throughout training, underscoring that stability is inherently architectural.
📝 Abstract
Despite powering modern AI, transformers remain mysteriously brittle to train. We develop a stability theory that explains why pre-LayerNorm works, why DeepNorm uses $N^{-1/4}$ scaling, and why warmup is necessary, all from first principles. Our framework has two pillars: (1) We derive the \emph{exact} operator norm of the softmax Jacobian, $\|J_{softmax}(u/τ)\|_{\infty\to 1} = θ(p)/τ$, where the balanced-mass factor $θ(p)\in[0,1]$ quantifies attention sensitivity. (2) We introduce a block-$\infty$/RMS geometry aligned with tokenwise computation, yielding Lipschitz bounds independent of sequence length. Using this framework, we prove that pre-LN preserves identity gradient paths while post-LN compounds LayerNorm Jacobians exponentially with depth, and we show that DeepNorm's $N^{-1/4}$ emerges from the quartic structure of attention's four projection matrices. We validate our theory on 774M-parameter models and find that, contrary to the intuition that attention sharpens during training to reduce sensitivity, $θ(p) \approx 1$ persists throughout. Transformer stability arises entirely from architectural gradient flow, not from attention dynamics. This finding changes how we reason about training: the architecture itself must handle sensitivity, not learned attention patterns.