🤖 AI Summary
This work investigates the out-of-sample prediction risk of high-dimensional unregularized least squares interpolation under proportional asymptotics, where both sample size and dimension grow at the same rate, with a focus on why overparameterized models generalize well under spiked covariance structures. By integrating random matrix theory, high-dimensional asymptotic analysis, and spectral methods, the study proposes a novel mechanism: prediction risk is governed by how signal energy distributes across the eigenspaces associated with covariance spikes. This mechanism reveals that geometric alignment between the regression coefficients and the spiked directions is key to benign overfitting. Under mild assumptions requiring only finite fourth moments, sharp asymptotic risk limits are established, offering a unified explanation for benign, moderate, or catastrophic overfitting. The analysis further elucidates how the number, strength, and structure of spikes jointly drive double descent phenomena, providing a cohesive theoretical framework for understanding the role of covariance structure in generalization.
📝 Abstract
This paper investigates the asymptotic behavior of the out-of-sample prediction risk of the high-dimensional ridgeless least-squares estimator when the feature dimension $p$ and the sample size $n$ grow proportionally. We consider a generalized spiked population covariance model with multiple latent factors, where the number of spiked eigenvalues may remain finite or increase with $n$, and the spiked eigenvalues may be bounded or diverge at arbitrary rates. Beyond characterizing the impact of covariance spectra, we reveal a new mechanism underlying benign overfitting: the prediction behavior of ridgeless interpolation is fundamentally governed by the alignment between the regression coefficient $\boldsymbolβ$ and the spiked eigenspaces of the population covariance matrix. In particular, we show that the signal energy distributed along latent spike directions determines whether interpolation leads to benign, tempered, or catastrophic overfitting. Our theoretical framework establishes sharp prediction risk limits under minimal moment conditions, requiring only finite fourth moments rather than Gaussianity. We characterize how the number, strength, and geometric structure of the spikes jointly influence the double-descent phenomenon. These results provide a unified understanding of when latent covariance structures facilitate or hinder generalization in overparameterized regression.