A Solvable Model of Adaptive Learning Rate Rescaling: Acceleration, Stability&Scaling

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates how effective learning rate drift induced by normalized updates affects training acceleration, stability, and resource scaling. Leveraging the random feature model and dynamical mean field theory (DMFT), the work elucidates the acceleration mechanisms and late-stage instabilities of normalized SGD, with theoretical predictions validated through linearized ResNet experiments. The primary contribution is the first solvable theoretical framework connecting normalization-induced acceleration, marginal stability, and width-batch allocation. Furthermore, this research quantifies the computational efficiency trade-offs between batch size and network width across distinct scaling regimes, empirically corroborating the predicted trends on CIFAR-5M.
📝 Abstract
A recurring design principle in modern optimizers is to decouple update magnitude from the raw gradient norm, yet its consequences for learning-curve and resource scaling remain unclear. We isolate this mechanism by studying normalized SGD in a random-feature model with power-law teacher and data covariance. Fixed-norm updates induce an effective learning rate that grows as gradients shrink. We derive a dynamical mean-field theory (DMFT) describing the joint dependence of the loss on training time, model width and batch size. Normalization initially accelerates SGD, mapping the power-law exponent $r_{\rm SGD}<1$ to $2r_{\rm SGD}/(1-r_{\rm SGD})$, with exponential convergence at $r_{\rm SGD}=1$ and formal finite-time convergence for $r_{\rm SGD}>1$. At finite step size, however, the same feedback ultimately breaks the acceleration and leads to marginal stability. The late-time theory yields width-limited, edge-of-stochastic-stability (EoSS), and deterministic edge-of-stability (EoS) regimes. These phases determine when larger batches or wider models reduce serial training time at comparable compute. We quantify in which of these phases increased batch size or width can compensate the excess compute use per step by fewer optimization steps to target loss. Linearized ResNet experiments on CIFAR-5M support the predicted acceleration, breakdown, and resource-scaling trends. Together, these results connect normalization-induced acceleration, EoS effects, and width--batch allocation within a solvable theory.
Problem

Research questions and friction points this paper is trying to address.

adaptive learning rate
gradient normalization
scaling laws
edge of stability
optimization dynamics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive Learning Rate Rescaling
Dynamical Mean-Field Theory
Normalized SGD
Edge of Stability
Resource Scaling
🔎 Similar Papers