Overfitting of Spectral Gradient Descent: How Matrix Geometry shapes Generalization and Implicit Bias

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates how the geometric structure—collapsed or dispersed—of shortcut signals in overparameterized matrix classification influences the generalization gap between spectral gradient descent (SpecGD) and standard gradient descent (GD) under noisy labels. Through low-rank signal decomposition, duality analysis, and simulations of exponential gradient dynamics, we reveal a second-order effect mechanism whereby merely altering shortcut geometry can reverse the relative performance of these algorithms. Our findings demonstrate that while single-step SpecGD generalizes well, prolonged training leads to significant degradation in directional generalization. This work elucidates the critical role of second-order optimization effects in shaping implicit bias, offering new theoretical insights into why certain optimization trajectories succeed or fail when learning from corrupted supervision in overparameterized regimes.
📝 Abstract
We study the generalization of spectral gradient descent (SpecGD) in overparameterized matrix classification with corrupted labels. Each input combines a shared low-rank signal with a rank-one sample-specific perturbation, referred to as a shortcut, that enables memorization but does not generalize. We contrast collapsed shortcuts, which share a singular direction, with dispersed shortcuts, which occupy distinct singular directions. Changing only this geometry can reverse the relative generalization of GD and SpecGD: collapsed shortcuts can favor SpecGD, while dispersed shortcuts can favor GD. In the dispersed regime, exact shortcut orthogonality eliminates the signal from the late-stage SpecGD direction, while vanishing random correlations collectively generate a small but generalization-relevant signal through a second-order effect. To identify the direction selected by SpecGD, which the spectral max-margin problem alone does not determine, we combine a refined analysis of its dual with the exponentiated-gradient dynamics of normalized loss weights. Finally, we show that a single SpecGD step can already interpolate and generalize well, while continued training converges to a direction with substantially worse generalization.
Problem

Research questions and friction points this paper is trying to address.

Spectral Gradient Descent
Overfitting
Generalization
Implicit Bias
Shortcut Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spectral Gradient Descent
Matrix Geometry
Implicit Bias
Shortcut Learning
Overparameterization