Score
Designs and fits linear or generalized linear predictive models with explicit penalty terms (L1, L2, Tikhonov) to control complexity, produce sparse (lasso) or shrunk (ridge) coefficient estimates, mitigate multicollinearity, and trade off bias and variance. Implements and analyzes estimators and solvers — deriving closed‑form ridge solutions, L1 optimization or emulation methods, scalable algorithms, and theoretical or empirical assessments (e.g., limiting behavior, stability, and baseline validation) for identification and prediction.
To address the substantial estimation bias and high prediction error of the Lasso in high-dimensional linear regression, this paper proposes a two-stage Lasso-Ridge refitting method: the first stage employs Lasso for variable selection, and the second stage applies Ridge regularization on the selected submodel to correct estimation bias. This constitutes the first systematic framework that simultaneously achieves variable selection consistency and prediction consistency. Theoretical analysis establishes an upper bound on the prediction error and demonstrates that the proposed method strictly dominates the standard Lasso—even under its optimal tuning rate. Monte Carlo simulations show that the method reduces prediction error by 15–30% across diverse high-dimensional settings, significantly improves estimation accuracy, and maintains robust variable selection and generalization performance under challenging scenarios such as low signal-to-noise ratios and highly correlated designs.
研究通过在Lasso等关联集上应用二次校正来改进预测,利用惩罚矩阵控制随机性并提供有限样本期望界限,以统一框架理解何时此类校正可提高预测。
This paper addresses the instability and interpretability degradation of parameter estimates in multivariate linear regression caused by multicollinearity. We propose a systematic mitigation framework based on generalized ridge regression. First, we derive closed-form analytical expressions for key diagnostic metrics—including variance, coefficient of variation, correlation coefficient, variance inflation factor (VIF), and condition number—under generalized ridge regression, thereby establishing a unified theoretical framework for quantifying and regulating multicollinearity in non-orthogonal design matrices. Through rigorous theoretical analysis and two numerical experiments, we demonstrate that the proposed method substantially improves estimation stability and model interpretability consistency. The framework provides an analytically tractable and empirically verifiable statistical tool for modeling high-dimensional collinear data.
This paper investigates the non-asymptotic statistical behavior of ridge regression in high-dimensional and even infinite-dimensional Hilbert spaces, moving beyond the classical proportional-scaling regime of random matrix theory. Methodologically, it integrates tools from random matrix theory, convex concentration inequalities, and spectral analysis to construct an equivalent diagonal sequence model. This enables, for the first time under non-proportional, non-asymptotic scaling, a multiplicative (1±Δ)-approximation of both bias and variance. Key contributions include: (1) dimension-free, explicit non-asymptotic upper bounds on the estimation risk; (2) exact risk characterization for spectrally regular covariates; (3) tight guarantees for benign overfitting in the overparameterized interpolation regime, achieving sharp control of generalization error under zero bias; and (4) unification and refinement of existing proportional asymptotic results.
This work addresses the problem of user-controllable fairness tuning in regression models. We propose a general fairness framework built upon ridge regression penalization. Its core innovation lies in explicitly incorporating fairness constraints into the ridge parameter selection process—achieving an adjustable trade-off between fairness and predictive performance via regularization with respect to sensitive attributes—and deriving partial closed-form solutions. The method supports multiple fairness definitions (e.g., demographic parity, equalized odds), extends to generalized linear models and kernelized settings, and corrects systematic experimental biases present in prior studies. Extensive experiments on six benchmark datasets demonstrate that, at comparable fairness levels, our approach significantly outperforms mainstream baselines—including Komiyama et al. and Zafar et al.—while simultaneously improving both goodness-of-fit and prediction accuracy.
This work proposes an iterative algorithm that automatically computes a near-optimal regularization strength for ridge regression under fixed design matrices, assuming limited samples, bounded covariance, and isotropic noise. The method leverages parameters from a generative model to construct, for the first time in the fixed-X setting, a provably convergent numerical procedure that achieves strong generalization to random designs through sample-based estimation. Experimental results demonstrate that, in both under- and over-parameterized regimes, the approach attains near-optimal generalization performance across varying sample sizes, dimensionality ratios, and noise levels, requiring only one or two additional ridge regression computations beyond the initial fit.
This paper addresses the lack of theoretical characterizations for the effective degrees of freedom (EDF) of adaptive Lasso and adaptive group Lasso. We propose, for the first time, unbiased EDF estimators under both orthogonal and non-orthogonal designs. Building upon the Stein’s unbiased risk estimation framework and leveraging structural analysis of penalized regression, we rigorously derive explicit closed-form EDF expressions, revealing the coupled influence of regularization parameters, coefficient signs, and least-squares initial estimates on model complexity. The proposed estimators require neither resampling nor approximation, offering both analytical tractability and broad applicability. Empirical evaluation on synthetic and real-world datasets demonstrates that our EDF estimator significantly improves model selection consistency of information criteria (e.g., AIC, BIC) and enhances the accuracy of prediction error estimation. This work provides a critical theoretical and practical tool for assessing and deploying adaptive regularization methods.
The rule of thumb regarding the relationship between the bias-variance tradeoff and model size plays a key role in classical machine learning, but is now well-known to break down in the overparameterized setting as per the double descent curve. In particular, minimum-norm interpolating estimators can perform well, suggesting the need for new tradeoff in these settings. Accordingly, we propose a regularization-sharpness tradeoff for overparameterized linear regression with an $\ell^p$ penalty. Inspired by the interpolating information criterion, our framework decomposes the selection penalty into a regularization term (quantifying the alignment of the regularizer and the interpolator) and a geometric sharpness term on the interpolating manifold (quantifying the effect of local perturbations), yielding a tradeoff analogous to bias-variance. Building on prior analyses that established this information criterion for ridge regularizers, this work first provides a general expression of the interpolating information criterion for $\ell^p$ regularizers where $p \ge 2$. Subsequently, we extend this to the LASSO interpolator with $\ell^1$ regularizer, which induces stronger sparsity. Empirical results on real-world datasets with random Fourier features and polynomials validate our theory, demonstrating how the tradeoff terms can distinguish performant linear interpolators from weaker ones.
In high-dimensional sparse linear regression, the choice of the regularization parameter critically affects the mean squared prediction error performance of the Lasso. This work, within a non-asymptotic framework, is the first to explicitly characterize the tuning threshold beyond which the Lasso estimator becomes inadmissible, revealing the pivotal roles of the design matrix and noise structure in determining this inadmissibility. To address this limitation, the authors propose a Lasso–Ridge hybrid method whose regularization path strictly dominates several classical Lasso tuning strategies. The proposed approach achieves substantially improved prediction accuracy, offering a theoretically grounded and practically superior criterion for regularization selection in high-dimensional regression settings.
This study addresses the challenge of characterizing the finite-sample distribution of ridge regression estimators, which hinders optimal regularization parameter selection and predictive performance. The authors propose a nonstandard asymptotic approach based on Gaussian approximation that accommodates heteroskedasticity and autocorrelation under a general data-generating mechanism. By introducing a local population parameter assumption and allowing the regularization parameter to vary with sample size, they establish the first effective finite-sample distributional approximation for low-dimensional ridge regression. Building on this approximation, they develop regularization parameter selection strategies that minimize either average or worst-case excess prediction risk, thereby substantially improving prediction accuracy.