๐ค AI Summary
This work addresses the neglect of input-space regularity mechanisms in theoretical analyses of deep neural network generalization error bounds, focusing specifically on the dynamic evolution of the empirical Lipschitz constant during the double-descent phenomenon. Methodologically, we conduct a systematic analysis of SGD training trajectories, gradient magnitude estimation, loss landscape curvature approximation, and phase segmentation of double descent. Our key contribution is the first empirical demonstration that the Lipschitz constant exhibits pronounced non-monotonic surge-and-decay behavior near the critical transition regimeโprecisely synchronized with peaks and troughs in test error. Furthermore, we establish that, near the critical point, the norm of parameter-space gradients tightly couples with the input-space Lipschitz constant; moreover, both model complexity and optimization dynamics are jointly characterized by loss curvature and the Euclidean distance of parameters from initialization. This work provides a novel geometric perspective and quantifiable mechanistic framework for understanding double descent.
๐ Abstract
Existing bounds on the generalization error of deep networks assume some form of smooth or bounded dependence on the input variable, falling short of investigating the mechanisms controlling such factors in practice. In this work, we present an extensive experimental study of the empirical Lipschitz constant of deep networks undergoing double descent, and highlight non-monotonic trends strongly correlating with the test error. Building a connection between parameter-space and input-space gradients for SGD around a critical point, we isolate two important factors -- namely loss landscape curvature and distance of parameters from initialization -- respectively controlling optimization dynamics around a critical point and bounding model function complexity, even beyond the training data. Our study presents novels insights on implicit regularization via overparameterization, and effective model complexity for networks trained in practice.