data-adaptive penalization

Design and analyze penalized estimators and penalty functions that are chosen adaptively from the observed data, including adaptive penalized regression and other data-driven penalty constructions. This work includes building shrinkage–threshold maps that induce equivalent penalty forms, deriving penalized-regression representations that recover ridge and lasso as limiting cases, and extending these constructions to handle correlated observations.

data-adaptivepenalization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.25
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$201K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

On the Degrees of Freedom of some Lasso procedures

Nov 26, 2025
MB
Mauro Bernardi
🏛️ University of Padova | University of Rome Tor Vergata

This paper addresses the lack of theoretical characterizations for the effective degrees of freedom (EDF) of adaptive Lasso and adaptive group Lasso. We propose, for the first time, unbiased EDF estimators under both orthogonal and non-orthogonal designs. Building upon the Stein’s unbiased risk estimation framework and leveraging structural analysis of penalized regression, we rigorously derive explicit closed-form EDF expressions, revealing the coupled influence of regularization parameters, coefficient signs, and least-squares initial estimates on model complexity. The proposed estimators require neither resampling nor approximation, offering both analytical tractability and broad applicability. Empirical evaluation on synthetic and real-world datasets demonstrates that our EDF estimator significantly improves model selection consistency of information criteria (e.g., AIC, BIC) and enhances the accuracy of prediction error estimation. This work provides a critical theoretical and practical tool for assessing and deploying adaptive regularization methods.

Enables accurate model selection and prediction error estimation.Estimates effective degrees of freedom for Adaptive Lasso and Group Lasso.Provides unbiased estimator for model complexity in adaptive regression.

This paper addresses the high mean squared error (MSE) and large variance of nonparametric estimators for causal inference under finite samples. We propose an asymptotically efficient and data-adaptive penalized shrinkage estimation framework. Our key contribution is the first derivation of the nonparametric efficiency bound for penalized target parameters, enabling the design of an L1/L2 joint tuning criterion that renders penalization a plug-in post-processing step for any asymptotically normal and efficient estimator. The method preserves asymptotic optimality while substantially reducing finite-sample MSE—simulations show average reductions of 20–40%. We apply it to quality evaluation of U.S. dialysis providers, achieving high-precision, interpretable estimation of heterogeneous treatment effects across multiple provider groups.

Applying penalization in causal inference for treatment effectsImproving finite-sample properties of non-parametric estimatorsReducing mean-squared error for large parameter sets

This study addresses the computational challenges posed by high-dimensional regularized estimating equations, which often exhibit non-gradient structures, asymmetric Jacobians, over-identification, non-smoothness, non-convexity, or nested optimization, rendering standard penalized methods inefficient. To tackle this, the paper proposes a unified formulation of such problems as fixed-point equations and systematically develops four computational paradigms—minimization-based, Dantzig-type, regularization-based, and fixed-point-based—integrating strategies from penalized optimization, constrained linear programming, iterative root-finding, and proximal fixed-point iterations. This cohesive framework substantially enhances both solvability and algorithmic stability for high-dimensional regularized estimating equations, demonstrating broad applicability to complex settings such as longitudinal data analysis and survival modeling.

computational challengesestimating equationshigh-dimensional statistics

A Unified Framework for Pattern Recovery in Penalized and Thresholded Estimation and its Geometry

Jul 19, 2023
PG
P. Graczyk
🏛️ Université d’Angers | TU Wien | Politechnika Wrocławska | Université Bourgogne Europe

This paper addresses the unified modeling of parameter pattern recovery in multi-class sparse regularized estimation (e.g., LASSO, SLOPE). Existing frameworks lack a general characterization of pattern structure and complexity across diverse polyhedral norms. Method: We propose a generalized pattern definition and complexity measure based on polyhedral norms and subdifferentials; extend LASSO’s irrepresentability condition to a geometric, noiseless recovery condition applicable to arbitrary polyhedral norms; and identify “accessibility”—a strictly weaker condition than irrepresentability—as sufficient for deterministic pattern recovery via thresholding estimators. Contributions: (i) A universal necessary and sufficient condition system for pattern recovery; (ii) Proof that the noiseless recovery condition serves as a unifying foundation for pattern identification across all polyhedral-regularized estimators; (iii) Demonstration that thresholding strategies substantially relax recovery requirements and enhance robustness, thereby unifying theory and enabling milder conditions.

Establishing thresholded estimators' pattern recovery under accessibilityExtending irrepresentability condition to broad estimator classesGeneralizing pattern recovery conditions for penalized estimators

Adapting to Misspecification

May 23, 2023
TB
Timothy B. Armstrong
🏛️ USC | UC Berkeley | NBER | UCL | CEMFI

This paper addresses the trade-off between robustness and efficiency under model misspecification, proposing an adaptive estimation framework that does not require a pre-specified upper bound on bias. The core challenge is to construct an estimator whose worst-case risk—relative to an oracle knowing the true bias bound—is minimized. Methodologically, we formulate an adaptive shrinkage estimator via weighted convex minimax optimization, calibrated against the oracle risk, and develop a lookup-table-based fast algorithm. Theoretically, our approach departs from conventional hypothesis-testing paradigms and achieves, for the first time, direct adaptation to the degree of misspecification. Empirically, the method substantially improves estimation accuracy and robustness across multiple canonical studies, offering both strong theoretical guarantees and practical computational efficiency.

Adapting to unknown bias bounds in restricted estimatorsBalancing robustness and efficiency in parameter estimationSolving weighted convex minimax for adaptive estimation

Latest Papers

What's happening recently
View more

This work addresses the lack of a unified framework for shrinkage, thresholding, and regularization methods in normal mean estimation by proposing a general class of estimators that encompasses both James–Stein-type and Lasso-type rules. Within this class, the authors derive the NOMAD estimator by minimizing a feasible, data-driven approximation to the risk. The approach is extended to settings with correlated observations and linear regression. A key innovation is the establishment of a unified risk minimization framework featuring a data-adaptive penalty, which—remarkably—achieves approximate risk consistency in correlated normal mean models for the first time. Theoretical analysis confirms that the estimator is consistent under both independent and correlated designs, admits an equivalent penalized regression formulation, and either subsumes or improves upon several classical methods in theory.

approximate risk minimizationnormal mean estimationquadratic risk

This study addresses the challenge of variable selection and parameter estimation in high-dimensional survival data with right-censoring and partially interval-censoring, particularly when covariates are highly correlated. The authors propose a linear rank regression approach for the accelerated failure time model that integrates Broken Adaptive Ridge (BAR) penalty with induced smoothing. This method, the first to incorporate BAR into a smoothed rank regression framework, enjoys the oracle property and grouping effect, accommodates multivariate partial interval censoring, and yields closed-form variance estimates. Computationally efficient implementation is achieved via a cyclic coordinate descent algorithm. Simulation studies demonstrate its superior performance over existing methods in both variable selection accuracy and estimation efficiency. The approach has been successfully applied to clinical datasets on primary biliary cirrhosis and colorectal cancer, and the accompanying R package aftPenCDA is publicly available on CRAN.

accelerated failure time modelcensored datacorrelated covariates

This study addresses the limited generalization performance of standard norm-based regularization in neural networks when dealing with high-dimensional or feature-correlated settings, where conventional methods inadequately control model complexity. To overcome this, the authors propose two covariance-aware adaptive regularization techniques: first, incorporating the input feature covariance structure into ℓ₂ weight decay to refine ridge-type penalties; second, combining ℓ₁ sparsity with covariance-informed ℓ₂ regularization to achieve structured sparsity. By innovatively embedding feature covariance information into classical Lasso and ridge regression frameworks, the approach enables more precise complexity control. Extensive experiments on Monte Carlo simulations and real-world datasets—including building cooling load prediction and leukemia cell classification—demonstrate that the proposed methods significantly outperform traditional regularization strategies and substantially enhance model generalization.

feature correlationhigh-dimensional datamodel complexity

This work addresses the challenge of efficiently estimating non-pathwise differentiable functionals—such as dose–response curves under continuous exposure—by proposing a novel approach that integrates higher-order highly adaptive lasso (HAL), spline basis projection, and Targeted Maximum Likelihood Estimation (TMLE). The method projects the target functional onto a finite-dimensional space spanned by higher-order spline bases to construct a pathwise differentiable approximation, and embeds LASSO regularization within the TMLE framework to enable fully data-adaptive inference without requiring pre-specified sieves or parametric models. The resulting estimator exhibits pointwise asymptotic normality, with convergence rates determined solely by the dimensionality and smoothness of the target functional, and demonstrates markedly superior performance over conventional HAL plug-in estimators in simulations.

continuous exposuredata-adaptive inferencedose-response curve

Hot Scholars

PO

Philipp Otto

University of Glasgow
Spatiotemporal statisticsenvironmetricsdata sciencenetwork modelling
JC

Jiguo Cao

Simon Fraser University
Functional Data AnalysisEstimating Differential EquationsMachine Learning
JB

Johannes Brachem

PhD Student, Georg-August-University Göttingen
Conditional Transformation ModelsBayesian StatisticsComputational StatisticsOpen Science
TK

Thomas Kneib

Chair of Statistics, Georg-August-University Göttingen
semiparametric regressionstatistical modelling