Grokking through the Lens of Minimum-Norm Interpolation

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear intrinsic mechanisms of delayed generalization (grokking) by quantifying the influence of inductive biases and signal structure. It constructs a statistical theoretical framework to analyze how regularization geometry and signal sparsity govern generalization during near-interpolation. Specifically, this work proposes a “0-1 generalization law,” revealing the statistical instability of minimum-norm interpolation and the boundaries of sparsity-induced gains. These theoretical findings are systematically validated through high-dimensional regression, convex norm analysis, and experiments on linear networks and Transformers. By precisely characterizing the evolution trajectories of training and generalization errors, this research demonstrates the significant advantages of sparse regularization under noiseless data conditions, thereby providing a rigorous theoretical foundation for understanding the grokking phenomenon.
📝 Abstract
Grokking shows that fitting the training data and learning the underlying signal can occur at very different stages. However, existing theories offer limited quantitative insight into how this delayed generalization depends on inductive bias and signal structure. Our work addresses the gap by developing a statistical theory that characterizes how regularization geometry and signal sparsity govern generalization near interpolation. In particular, we focus on the prototypical setting of high-dimensional regression and identify regimes in which sparsity-promoting regularization makes exact interpolation much more accurate than approximate fitting. In strongly overparameterized noiseless problems, we prove a zero--one generalization law and construct a family of convex norms whose interpolators transition from the trivial risk of the all-zero predictor to exact recovery, while keeping the training error equal to $0$. Furthermore, when feature dimension and sample size are proportional, we provide a precise characterization of training and generalization errors along $\ell_r$-regularization paths. This in turn allows us to quantify the generalization gain that remains near interpolation: we show that this gain increases as the norm becomes more sparsity-promoting and as the target becomes sparser, with a sharp drop in generalization reached for noiseless data and $\ell_1$ regularization. Experiments on diagonal linear networks and transformers trained on modular arithmetic demonstrate the generality of our theoretical predictions. Finally, beyond grokking, our work reveals a statistical instability in minimum-norm interpolation: small perturbations in the regularization strength can lead to drastically different generalization, while preserving small training error.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Grokking
Minimum-Norm Interpolation
Sparsity-Promoting Regularization
High-Dimensional Regression
Statistical Instability
💼 Related Jobs
No related jobs found.
Gil Kur
Gil Kur
Postdoc at ETH Zürich
Nonparametric statisticshigh dimensional statisticsconvex geometrylearning theory
I
Ileana Rugina
Institute of Science and Technology Austria
C
Clémentine Carla Juliette Dominé
Institute of Science and Technology Austria, Harvard University
Marco Mondelli
Marco Mondelli
Professor, IST Austria
Machine LearningData ScienceCoding TheoryInformation theory