online ridge regression

Designs and implements algorithms that incrementally update linear regression parameter estimates using L2 (ridge) regularization, producing recursive least-squares-style update rules that can be applied to streaming data. Analyzes and derives recursive inference equations, implements efficient online solvers, and optionally maintains online estimates of parameter uncertainty.

onlineridgeregression

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.45
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work proposes an iterative algorithm that automatically computes a near-optimal regularization strength for ridge regression under fixed design matrices, assuming limited samples, bounded covariance, and isotropic noise. The method leverages parameters from a generative model to construct, for the first time in the fixed-X setting, a provably convergent numerical procedure that achieves strong generalization to random designs through sample-based estimation. Experimental results demonstrate that, in both under- and over-parameterized regimes, the approach attains near-optimal generalization performance across varying sample sizes, dimensionality ratios, and noise levels, requiring only one or two additional ridge regression computations beyond the initial fit.

generalizationisotropic noiselinear regression

This work addresses the inconsistency among multi-granularity forecasts in online hierarchical time series prediction by proposing a reconciliation method that explicitly models hierarchical relationships through a graph structure. The approach characterizes forecast residuals using a matrix normal distribution and formulates a multivariate linear regression framework, integrating ridge regression, Bayesian estimation, and shrinkage principles. An efficient online recursive inference mechanism is developed to enable adaptive forecast reconciliation and uncertainty quantification. The method is validated on a district heating load forecasting task, demonstrating its effectiveness. To support practical deployment, the authors release PyOnlineForecast, an open-source toolkit for online hierarchical forecasting.

forecast coordinationhierarchical reconciliationlinear models

In high-dimensional linear regression with strongly anisotropic covariates and true coefficients aligned with top eigendirections of the covariance matrix, conventional ℓ₂-shrinkage estimators (e.g., ridge regression) can underperform the minimum-ℓ₂-norm interpolator. This work demonstrates that *moderate amplification*—scaling the minimum-norm interpolator by a constant greater than one—substantially reduces generalization error, challenging the long-held belief that shrinkage universally dominates interpolation. We establish the first tight upper and lower bounds on the generalization error under diverging aspect ratios (n/d → 0) for anisotropic covariance structures, proving that amplified interpolation achieves the optimal statistical rate. Our theoretical analysis, combining data splitting and Gaussian random projection techniques, rigorously characterizes the mechanism and limits of amplification. Empirical results confirm that the amplified interpolator consistently outperforms classical regularized estimators on real-world anisotropic data.

Addressing anisotropic covariances in high-dimensional linear regressionImproving generalization without traditional shrinkage regularizationReducing test error by inflating minimum norm interpolators

Existing pedagogical treatments of linear regression often lack rigorous theoretical unification across geometric, algebraic, and statistical perspectives, and seldom establish formal optimality guarantees or bridge frequentist and Bayesian interpretations. Method: This paper constructs a rigorous theoretical framework for linear regression targeting readers familiar with ordinary least squares (OLS), integrating geometric projection, matrix algebra, and statistical inference. It introduces Gaussian noise assumptions to derive maximum likelihood estimation, exact sampling distributions (t- and F-tests, confidence intervals), and Bayesian linear regression with closed-form posterior inference. Contribution/Results: We provide the first rigorous proof that OLS achieves the Cramér–Rao lower bound among unbiased linear estimators—establishing its theoretical optimality. The work unifies frequentist and Bayesian paradigms, yielding analytically tractable inference tools. The resulting verifiable statistical pipeline serves as an interpretable linear foundation and error-analysis benchmark for nonlinear models, including deep learning.

Analyze distribution theory and unbiased estimators in linear modelsExplore properties of least squares approximation in linear modelsIntroduce rigorous theories behind linear models for regression

Improving Linear System Solvers for Hyperparameter Optimisation in Iterative Gaussian Processes

May 28, 2024
JA
Jihao Andreas Lin
🏛️ University of Cambridge | MPI for Intelligent Systems

In large-scale Gaussian process (GP) hyperparameter optimization, iterative linear solvers—such as conjugate gradient (CG)—induce inefficiency in computing gradients of the marginal likelihood due to repeated, costly matrix-vector operations. Method: We propose a general-purpose optimization framework integrating pathwise gradient estimation, solver warm-starting, and budget-aware early stopping. The framework is agnostic to the underlying iterative solver and supports CG, alternating projections, and stochastic gradient descent. Contribution/Results: Our approach substantially alleviates the accuracy–efficiency trade-off in gradient estimation. Experiments demonstrate up to 72× speedup over standard CG when solving to full convergence. Under early stopping, the average residual norm drops to one-seventh of that achieved by baseline methods, significantly shortening hyperparameter optimization time while preserving convergence stability and gradient estimation accuracy.

Gaussian ProcessesHyperparameter OptimizationLarge-scale Datasets

Latest Papers

What's happening recently
View more

This work investigates the “grokking” phenomenon in over-parameterized linear models trained via gradient descent—wherein models suddenly transition from prolonged overfitting to perfect generalization. Within a ridge regression framework augmented with weight decay, the authors systematically identify three distinct phases: initial overfitting, an extended period of poor generalization, and eventual convergence of generalization error to zero. Their theoretical analysis establishes, for the first time, a precise quantitative relationship between the onset time of grokking and training hyperparameters, demonstrating that grokking can be controlled or even eliminated through careful hyperparameter tuning. Experiments further confirm that this mechanism persists in nonlinear neural networks, indicating that grokking arises from optimization dynamics rather than architectural deficiencies.

generalization delaygradient descentgrokking

This study investigates the statistical properties of Lagrange multipliers in constrained maximum likelihood estimation and least squares problems, along with their implications for numerical optimization. Leveraging large-sample theory, it establishes that under correctly specified models, Lagrange multipliers converge in probability to zero as the sample size grows, a result extended to high-dimensional settings such as deep learning. Building on this asymptotic behavior, the work provides the first statistical justification for initializing Lagrange multipliers at zero and integrates this insight into constrained optimization algorithms, including augmented Lagrangian methods and sequential quadratic programming. Numerical experiments demonstrate that this initialization strategy substantially enhances algorithmic stability and convergence efficiency in applications such as constrained regression and dynamic discrete choice models.

asymptotic behaviorconstrained optimizationLagrange multipliers

This work addresses the efficient computation of L1 regularization paths in linear models, encompassing applications such as LASSO, linear support vector machines, and L1-regularized Kalman smoothing. The authors propose a factor graph approach based on parametric Gaussian message passing, which employs forward–backward recursions to separately handle L1 penalties on predictors and responses, yielding a pair of dual algorithms. This is the first method to integrate parametric Gaussian message passing into L1 path computation, substantially extending sparse modeling capabilities within a state-space framework. The algorithm is highly general, relying primarily on matrix multiplications, and achieves computational complexity that improves upon existing approaches in certain regimes.

Kalman smoothingL1 regularizationLASSO

Hot Scholars

HZ

Huiping Zhuang

Associate Professor, South China University of Technology
Continual LearningMulti-ModalEmbodied AILarge Model
KW

Kaizheng Wang

Columbia University
Machine LearningStatisticsOptimization
LX

Liyuan Xu

Secondmind
Machine LearningOnline Learning
RH

Run He

South China university of Technology
Deep LearningContinual LearningFederated LearningLLM
DF

Di Fang

South China University of Technology
Continual Learning