regression

Designs, fits, and evaluates statistical and machine‑learning models that estimate a continuous outcome as a function of one or more predictor variables. This includes selecting appropriate linear or nonlinear regression formulations, applying regularization and feature selection, performing diagnostics and assumption checks, quantifying uncertainty in estimates, and comparing model performance.

regression

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.9
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$206K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the challenge of variable selection in linear regression by proposing a data-driven approach based on artificial neural networks (ANNs). The method leverages statistical quantities derived from ordinary least squares (OLS) estimation to automatically assess variable significance, marking the first end-to-end application of ANNs within the OLS framework for variable selection. This yields a scalable and intelligent modeling paradigm. Empirical evaluations demonstrate that the proposed approach consistently achieves superior variable selection accuracy across diverse sample sizes and error variance configurations, outperforming conventional techniques such as forward/backward selection, AIC, BIC, and LASSO. Its practical utility is further validated on the WHO life expectancy dataset. The authors publicly release a pretrained ANN model capable of handling up to 100 predictors.

Artificial IntelligenceLinear ModelsVariable Selection

Fundamentals of Regression

Nov 27, 2025
MA
Miguel A. Mendez
🏛️ von Karman Institute for Fluid Dynamics

This paper addresses the limited physical interpretability and poor generalizability of conventional data-driven regression methods. To this end, we propose a physics-informed regression modeling paradigm that systematically embeds first-principles constraints, differential equation priors, and conservation laws into statistical regression, curve fitting, and supervised learning frameworks—thereby enabling deep integration of machine learning with classical numerical methods (e.g., finite differences, spectral methods). Our key contributions are threefold: (i) we establish a systematic taxonomy tracing the evolution of regression from purely statistical to physics-guided formulations; (ii) we construct a theoretical bridge linking computational science and scientific machine learning; and (iii) the proposed framework significantly enhances model extrapolation capability, robustness, and physical consistency. As a result, it offers an interpretable, verifiable, and cross-disciplinary modeling paradigm for complex scientific and engineering problems.

Combines machine learning and numerical methods for physicsIt integrates physical knowledge with machine learning methodsRegression moves from data-driven to physics-informed formalism

Conformalized Regression for Continuous Bounded Outcomes

Jul 18, 2025
ZW
Zhanli Wu
🏛️ King’s College London | University College London

Existing regression methods for bounded continuous response variables (e.g., proportions, ratios) typically yield only point estimates or asymptotically valid prediction intervals, failing to simultaneously guarantee finite-sample coverage accuracy and robustness to model misspecification. Method: We propose a novel conformal prediction framework integrating transformation models with Bayesian regression. It introduces a heteroscedasticity-adapted nonconformity measure, unifies full and split conformal inference strategies, and employs residual transformation to enhance robustness. Contribution/Results: We establish theoretical guarantees of both marginal and conditional coverage validity. Monte Carlo simulations and empirical studies demonstrate that the method achieves exact nominal coverage even in small samples—substantially outperforming alternatives such as the bootstrap. The framework is particularly effective for bounded outcomes, offering improved calibration, robustness to distributional assumptions, and computational feasibility without sacrificing statistical rigor.

Addressing heteroscedasticity in regression with bounded outcomesEnsuring valid predictive coverage under model misspecificationPredicting bounded continuous outcomes accurately

Estimating and evaluating counterfactual prediction models

Aug 24, 2023
CB
Christopher B. Boyer
🏛️ Cleveland Clinic Research | Case Western Reserve University | Harvard T.H. Chan School of Public Health | Richard A. and Susan F. Smith Center for Outcomes Research | Beth Israel Deaconess Medical Center | Brown University School of Public Health

Counterfactual prediction under evolving intervention policies or hypothetical decision scenarios remains challenging due to unobservable potential outcomes, hindering model identifiability, evaluation, and generalization. Method: We propose the first systematic theoretical framework addressing this challenge—comprising (i) identifiability conditions for counterfactual prediction models, (ii) a performance evaluation system targeting loss, AUC, and calibration, and (iii) robust hyperparameter selection under model misspecification. Our approach integrates causal inference principles, doubly robust estimation, and loss-driven evaluation metric design. Contribution/Results: Validated via simulation studies and a real-world clinical application—cardiovascular risk prediction in statin-naïve populations—the framework significantly improves out-of-distribution generalization and clinical decision reliability in counterfactual settings.

Estimating counterfactual prediction models under different treatment policiesEvaluating model performance without observed potential outcomesProviding valid performance estimates under model misspecification

kNN Algorithm for Conditional Mean and Variance Estimation with Automated Uncertainty Quantification and Variable Selection

Feb 02, 2024
MM
Marcos Matabuena
🏛️ Harvard T.H. Chan School of Public Health | University of Santiago de Compostela | University of California, Los Angeles

This paper addresses the challenge of conditional distribution reconstruction under high-dimensional covariates. We propose a novel k-nearest neighbors (kNN) semi-parametric regression method that jointly estimates the conditional mean and variance to reconstruct the conditional density function and quantify predictive uncertainty. Our key contributions are: (1) the first semi-parametric ROC curve estimation framework within kNN; (2) a theoretically grounded, adaptive k-selection algorithm; and (3) a conditional distribution modeling framework integrated with variable selection. Under low-dimensional structural assumptions—such as intrinsic dimensionality or sparsity—we establish consistency and derive the optimal nonparametric convergence rate. Simulation studies demonstrate substantial improvements over conventional kNN methods in estimation accuracy and uncertainty calibration. Furthermore, two real-world biomedical applications confirm the method’s robustness and practical utility in high-dimensional settings.

Automating variance selection to improve empirical performanceEstimating conditional mean and variance via k-NN regressionReconstructing conditional distributions for generative models

Latest Papers

What's happening recently
View more

This work proposes a unified inference framework based on parametric programming to address the bias in statistical inference for regression coefficients following Lasso variable selection in generalized linear models. By locally linearizing the maximum likelihood estimator, the method constructs a linear relationship between pseudo-responses and covariates, thereby extending parametric programming—previously limited to Gaussian settings—to non-Gaussian response distributions. This extension enables valid post-selection inference across a range of exponential family models, including logistic and Poisson regression. The approach maintains computational efficiency while substantially improving inferential accuracy. Simulation studies demonstrate that, in non-Gaussian settings, the proposed method effectively corrects the naive inference that ignores the selection process and achieves higher statistical efficiency compared to existing approaches such as the polyhedral method.

generalized linear modelsLassopost-selection inference

This work proposes an additive nonlinear Bayesian regression framework to address the challenge of functional outputs in complex computer simulations that are jointly influenced by functional predictors defined over a fixed spatial domain and global scalar variables varying across simulation runs. The approach introduces a novel functional Gaussian process (fGP) prior that simultaneously models the unknown nonlinear effect of global variables across the entire spatial domain and captures spatially varying coefficients of local functional predictors through Gaussian processes. By explicitly encoding spatial dependence in the global effects, the fGP enables an interpretable decomposition of contributions from these two multiscale predictor types and provides principled uncertainty quantification. Experiments on synthetic data and the SLOSH hurricane storm surge model demonstrate that the method achieves high predictive accuracy alongside reliable uncertainty estimates.

function-on-function regressionfunctional datamulti-scale predictors

This study addresses selection bias induced by covariates in the estimation of continuous treatment effects by proposing a two-stage kernel ridge regression approach. In the first stage, the method models the joint dependence of the outcome on both the treatment variable and covariates; in the second stage, it constructs pseudo-outcomes that correct for distributional shifts to estimate the average treatment effect. The core innovation lies in an estimator that adaptively exploits potential structural simplifications in the treatment effect function, coupled with a fully data-driven model selection procedure that requires no prior knowledge. This framework simultaneously adapts to unknown degrees of overlap and kernel eigenvalue decay rates, achieving theoretical guarantees for adaptivity to both overlap conditions and function smoothness, thereby yielding more accurate estimates of continuous treatment effects.

causal inferenceconfoundingcontinuous treatment effects

This study addresses the challenges of tuning and evaluation ambiguity in semiparametric and high-dimensional models arising from redundant components introduced by kernel smoothing or basis expansions. The authors propose an induced replication framework that leverages the principle of parametric inference, transforming model assessment into an in-sample prediction error problem by exploiting known replication mechanisms embedded within the model. Building upon Fisher’s concepts of sufficiency and conditional sufficiency, this approach replaces conventional out-of-sample prediction and is applicable to proportional hazards models, time-varying Poisson processes, and confidence set construction for sparse regression. Both theoretical analysis and numerical experiments demonstrate that the method accurately controls nominal error rates under correct model specification and exhibits high sensitivity to semiparametric misspecification.

induced replicationmodel assessmentnuisance parameters

This study addresses the challenge that data-driven variable selection methods—such as Lasso and its adaptive variants—undermine the validity of classical statistical inference in Cox proportional hazards models, particularly leading to inflated false positive rates in right-censored survival data. The authors systematically evaluate the post-selection inference performance of sample splitting, exact post-selection inference, and debiased Lasso approaches. For the first time, these methods are comprehensively compared within a simulation framework designed to closely mimic real-world biomedical scenarios, and their practical utility is further validated using publicly available datasets. The findings elucidate the trade-offs among these methods in controlling Type I error rates and achieving estimation accuracy, thereby offering reliable and practical inference strategies for high-dimensional survival data analysis.

Cox modelpost-selection inferenceright-censoring