regularized regression

Designs and fits linear or generalized linear predictive models with explicit penalty terms (L1, L2, Tikhonov) to control complexity, produce sparse (lasso) or shrunk (ridge) coefficient estimates, mitigate multicollinearity, and trade off bias and variance. Implements and analyzes estimators and solvers — deriving closed‑form ridge solutions, L1 optimization or emulation methods, scalable algorithms, and theoretical or empirical assessments (e.g., limiting behavior, stability, and baseline validation) for identification and prediction.

regularizedregression

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.29
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$204K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

To address the substantial estimation bias and high prediction error of the Lasso in high-dimensional linear regression, this paper proposes a two-stage Lasso-Ridge refitting method: the first stage employs Lasso for variable selection, and the second stage applies Ridge regularization on the selected submodel to correct estimation bias. This constitutes the first systematic framework that simultaneously achieves variable selection consistency and prediction consistency. Theoretical analysis establishes an upper bound on the prediction error and demonstrates that the proposed method strictly dominates the standard Lasso—even under its optimal tuning rate. Monte Carlo simulations show that the method reduces prediction error by 15–30% across diverse high-dimensional settings, significantly improves estimation accuracy, and maintains robust variable selection and generalization performance under challenging scenarios such as low signal-to-noise ratios and highly correlated designs.

Improves prediction performance in high-dimensional linear regressionReduces Lasso's estimation bias and prediction errorRetains Lasso's advantages like variable selection consistency

Generalized Ridge Regression: Applications to Nonorthogonal Linear Regression Models

Apr 08, 2025
RS
Rom'an Salmer'on G'omez
🏛️ University of Granada

This paper addresses the instability and interpretability degradation of parameter estimates in multivariate linear regression caused by multicollinearity. We propose a systematic mitigation framework based on generalized ridge regression. First, we derive closed-form analytical expressions for key diagnostic metrics—including variance, coefficient of variation, correlation coefficient, variance inflation factor (VIF), and condition number—under generalized ridge regression, thereby establishing a unified theoretical framework for quantifying and regulating multicollinearity in non-orthogonal design matrices. Through rigorous theoretical analysis and two numerical experiments, we demonstrate that the proposed method substantially improves estimation stability and model interpretability consistency. The framework provides an analytically tractable and empirically verifiable statistical tool for modeling high-dimensional collinear data.

Analyze variance and correlation coefficientsIllustrate results with numerical examplesMitigate multicollinearity in linear regression models

Dimension free ridge regression

Oct 16, 2022
CC
Chen Cheng
🏛️ Stanford University | Institute for Advanced Studies | Princeton

This paper investigates the non-asymptotic statistical behavior of ridge regression in high-dimensional and even infinite-dimensional Hilbert spaces, moving beyond the classical proportional-scaling regime of random matrix theory. Methodologically, it integrates tools from random matrix theory, convex concentration inequalities, and spectral analysis to construct an equivalent diagonal sequence model. This enables, for the first time under non-proportional, non-asymptotic scaling, a multiplicative (1±Δ)-approximation of both bias and variance. Key contributions include: (1) dimension-free, explicit non-asymptotic upper bounds on the estimation risk; (2) exact risk characterization for spectrally regular covariates; (3) tight guarantees for benign overfitting in the overparameterized interpolation regime, achieving sharp control of generalization error under zero bias; and (4) unification and refinement of existing proportional asymptotic results.

Characterizes ridge regression for high-dimensional Hilbert covariatesEstablishes non-asymptotic bounds for bias and varianceExtends ridge regression analysis beyond proportional asymptotics

Achieving fairness with a simple ridge penalty

May 18, 2021
MS
M. Scutari
🏛️ Istituto Dalle Molle di Studi sull’Intelligenza Artificiale (IDSIA) | London School of Economics | University of Oxford | UBS

This work addresses the problem of user-controllable fairness tuning in regression models. We propose a general fairness framework built upon ridge regression penalization. Its core innovation lies in explicitly incorporating fairness constraints into the ridge parameter selection process—achieving an adjustable trade-off between fairness and predictive performance via regularization with respect to sensitive attributes—and deriving partial closed-form solutions. The method supports multiple fairness definitions (e.g., demographic parity, equalized odds), extends to generalized linear models and kernelized settings, and corrects systematic experimental biases present in prior studies. Extensive experiments on six benchmark datasets demonstrate that, at comparable fairness levels, our approach significantly outperforms mainstream baselines—including Komiyama et al. and Zafar et al.—while simultaneously improving both goodness-of-fit and prediction accuracy.

Controlling sensitive attribute effects using ridge penalty selectionEstimating regression models with user-defined fairness levelsExtending fair regression to various models and fairness definitions

Latest Papers

What's happening recently
View more

This work proposes an iterative algorithm that automatically computes a near-optimal regularization strength for ridge regression under fixed design matrices, assuming limited samples, bounded covariance, and isotropic noise. The method leverages parameters from a generative model to construct, for the first time in the fixed-X setting, a provably convergent numerical procedure that achieves strong generalization to random designs through sample-based estimation. Experimental results demonstrate that, in both under- and over-parameterized regimes, the approach attains near-optimal generalization performance across varying sample sizes, dimensionality ratios, and noise levels, requiring only one or two additional ridge regression computations beyond the initial fit.

generalizationisotropic noiselinear regression

On the Degrees of Freedom of some Lasso procedures

Nov 26, 2025
MB
Mauro Bernardi
🏛️ University of Padova | University of Rome Tor Vergata

This paper addresses the lack of theoretical characterizations for the effective degrees of freedom (EDF) of adaptive Lasso and adaptive group Lasso. We propose, for the first time, unbiased EDF estimators under both orthogonal and non-orthogonal designs. Building upon the Stein’s unbiased risk estimation framework and leveraging structural analysis of penalized regression, we rigorously derive explicit closed-form EDF expressions, revealing the coupled influence of regularization parameters, coefficient signs, and least-squares initial estimates on model complexity. The proposed estimators require neither resampling nor approximation, offering both analytical tractability and broad applicability. Empirical evaluation on synthetic and real-world datasets demonstrates that our EDF estimator significantly improves model selection consistency of information criteria (e.g., AIC, BIC) and enhances the accuracy of prediction error estimation. This work provides a critical theoretical and practical tool for assessing and deploying adaptive regularization methods.

Enables accurate model selection and prediction error estimation.Estimates effective degrees of freedom for Adaptive Lasso and Group Lasso.Provides unbiased estimator for model complexity in adaptive regression.

The rule of thumb regarding the relationship between the bias-variance tradeoff and model size plays a key role in classical machine learning, but is now well-known to break down in the overparameterized setting as per the double descent curve. In particular, minimum-norm interpolating estimators can perform well, suggesting the need for new tradeoff in these settings. Accordingly, we propose a regularization-sharpness tradeoff for overparameterized linear regression with an $\ell^p$ penalty. Inspired by the interpolating information criterion, our framework decomposes the selection penalty into a regularization term (quantifying the alignment of the regularizer and the interpolator) and a geometric sharpness term on the interpolating manifold (quantifying the effect of local perturbations), yielding a tradeoff analogous to bias-variance. Building on prior analyses that established this information criterion for ridge regularizers, this work first provides a general expression of the interpolating information criterion for $\ell^p$ regularizers where $p \ge 2$. Subsequently, we extend this to the LASSO interpolator with $\ell^1$ regularizer, which induces stronger sparsity. Empirical results on real-world datasets with random Fourier features and polynomials validate our theory, demonstrating how the tradeoff terms can distinguish performant linear interpolators from weaker ones.

generalizationlinear interpolationoverparameterization

In high-dimensional sparse linear regression, the choice of the regularization parameter critically affects the mean squared prediction error performance of the Lasso. This work, within a non-asymptotic framework, is the first to explicitly characterize the tuning threshold beyond which the Lasso estimator becomes inadmissible, revealing the pivotal roles of the design matrix and noise structure in determining this inadmissibility. To address this limitation, the authors propose a Lasso–Ridge hybrid method whose regularization path strictly dominates several classical Lasso tuning strategies. The proposed approach achieves substantially improved prediction accuracy, offering a theoretically grounded and practically superior criterion for regularization selection in high-dimensional regression settings.

high-dimensional regressioninadmissibilityLasso

This study addresses the challenge of characterizing the finite-sample distribution of ridge regression estimators, which hinders optimal regularization parameter selection and predictive performance. The authors propose a nonstandard asymptotic approach based on Gaussian approximation that accommodates heteroskedasticity and autocorrelation under a general data-generating mechanism. By introducing a local population parameter assumption and allowing the regularization parameter to vary with sample size, they establish the first effective finite-sample distributional approximation for low-dimensional ridge regression. Building on this approximation, they develop regularization parameter selection strategies that minimize either average or worst-case excess prediction risk, thereby substantially improving prediction accuracy.

bias-variance tradeofffinite-sample distributionprediction risk

Hot Scholars

MM

Marco Mondelli

Professor, IST Austria
Machine LearningData ScienceCoding TheoryInformation theory
MK

Masahiro Kato

Mizuho-DL Financial Technology Co., Ltd. / The University of Tokyo
Economics
OD

Oliver Dukes

Ghent University
BiostatisticsCausal inference
LL

Lin Liu

University of Science and Technology of China
image/video generationcomputer visiondeep learningcomputational photography