🤖 AI Summary
This work investigates the mechanism underlying grokking—the phenomenon where test performance lags significantly behind training loss convergence. By analyzing linear models optimized via heavy-ball momentum with weight decay, the study introduces the notion of a “grokking subspace” and demonstrates that weight decay alone acts as a restorative force driving slow dissipative relaxation. The authors derive precise discrete- and continuous-time relaxation laws and theoretically distinguish between coupled L2 regularization and decoupled weight decay, yielding testable causal predictions. All theoretical claims are verified without fitted parameters in analytically tractable synthetic models. Furthermore, on modular addition tasks, delayed grokking is observed to obey the scaling law $(1-\beta)/(\eta\lambda)$, with late-stage relaxation aligning closely with the theoretically predicted timescale.
📝 Abstract
Delayed generalization, or grokking, remains poorly understood despite extensive empirical study. We identify an exactly solvable late-time relaxation mechanism for grokking in linear models trained with full-batch heavy-ball optimization and weight decay, together with a locally quadratic extension to nonlinear neural networks. Our analysis reveals a distinguished population-active component of the empirical null space, which we call the grokking subspace. Along this subspace, the training predictions remain unchanged, leaving weight decay as the sole restoring force and giving rise to a slow dissipative relaxation governed by an exact discrete-time and continuous-time law. We show that only this subspace contributes to the slow asymptotic decay of the population risk and derive explicit iteration-scale predictions for the grokking time, recovering the familiar $(1-β)/(ηλ)$ scaling in the weak-regularization regime. The theory further predicts distinct effects of optimizer choice, distinguishing coupled $L_2$ regularization from decoupled weight decay, and yields causal predictions for interventions that modify the grokking component. We verify all theoretical identities without fitted parameters in a synthetic model where every subspace and relaxation rate is computable in closed form. We further observe genuine delayed generalization in modular addition, where the measured delay follows the predicted scaling and the late-time relaxation agrees closely with the theoretical clock.