🤖 AI Summary
In high-dimensional linear regression with strongly anisotropic covariates and true coefficients aligned with top eigendirections of the covariance matrix, conventional ℓ₂-shrinkage estimators (e.g., ridge regression) can underperform the minimum-ℓ₂-norm interpolator. This work demonstrates that *moderate amplification*—scaling the minimum-norm interpolator by a constant greater than one—substantially reduces generalization error, challenging the long-held belief that shrinkage universally dominates interpolation. We establish the first tight upper and lower bounds on the generalization error under diverging aspect ratios (n/d → 0) for anisotropic covariance structures, proving that amplified interpolation achieves the optimal statistical rate. Our theoretical analysis, combining data splitting and Gaussian random projection techniques, rigorously characterizes the mechanism and limits of amplification. Empirical results confirm that the amplified interpolator consistently outperforms classical regularized estimators on real-world anisotropic data.
📝 Abstract
Hastie et al. (2022) found that ridge regularization is essential in high dimensional linear regression $y=β^Tx + ε$ with isotropic co-variates $xin mathbb{R}^d$ and $n$ samples at fixed $d/n$. However, Hastie et al. (2022) also notes that when the co-variates are anisotropic and $β$ is aligned with the top eigenvalues of population covariance, the "situation is qualitatively different." In the present article, we make precise this observation for linear regression with highly anisotropic covariances and diverging $d/n$. We find that simply scaling up (or inflating) the minimum $ell_2$ norm interpolator by a constant greater than one can improve the generalization error. This is in sharp contrast to traditional regularization/shrinkage prescriptions. Moreover, we use a data-splitting technique to produce consistent estimators that achieve generalization error comparable to that of the optimally inflated minimum-norm interpolator. Our proof relies on apparently novel matching upper and lower bounds for expectations of Gaussian random projections for a general class of anisotropic covariance matrices when $d/n o infty$.