🤖 AI Summary
This work uncovers the mechanism by which the interaction between normalization and weight decay during deep neural network training triggers abrupt loss spikes: the scale invariance induced by normalization causes weight decay to continuously shrink the weight norm, leading to a sharp increase in loss landscape sharpness and optimization instability. To address this, the study introduces a novel concept—“weight norm criticality”—which elucidates why excessive weight decay, while improving generalization, inherently destabilizes training, and provides testable theoretical predictions. Through rigorous theoretical analysis and extensive experiments across multiple architectures, the authors establish a causal relationship between weight norm evolution and loss spikes, offering a principled foundation for balancing regularization strength and training stability.
📝 Abstract
Most explanations of training instability focus on \emph{learning-rate criticality}, typically characterized by the Edge of Stability, beyond which optimization becomes unstable. We argue that, in practical deep neural network training, there is an additional and often overlooked \emph{weight-norm criticality}. This criticality is induced by the interaction between normalization (which introduces scale-invariant components) and weight decay (which persistently shrinks parameter norms). As the weight decay coefficient increases, the norms of scale-invariant weights are progressively driven toward zero. Meanwhile, the sharpness of the loss landscape increases rapidly, destabilizing the optimization dynamics and resulting in abrupt loss spikes. This perspective provides a rationale for why weight penalties can improve generalization yet cannot be made arbitrarily strong: excessive decay drives scale-invariant weight norms past a critical boundary and destabilizes training. Our work provides a new mechanistic understanding of loss spikes through the lens of \emph{weight-norm criticality}. Moreover, \emph{weight-norm criticality} yields testable predictions that we validate empirically in networks with scale-invariant components, providing empirical support for the proposed mechanism.