regression modeling

Building and evaluating statistical regression and classification models (including parametric and mechanistic variants) to predict continuous and categorical outcomes, adapt calibration tools, and quantify sources of variance. This includes selecting features, modeling hierarchical observation processes, and validating predictive performance and calibration.

regressionmodeling

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the unification of calibration concepts across classification and regression tasks, aiming to ensure consistency between predicted distributions and observed outcomes for diverse data types—continuous, discrete, nominal, and binary. The work introduces modal calibration for nominal outcomes and establishes a hierarchical framework distinguishing full, partial, and average calibration. It proposes a generalized definition of calibration based on predictive distribution functionals—such as means, quantiles, and event probabilities—and leverages probability integral transforms alongside constructive algorithms for analysis. Key contributions include demonstrating the logical independence between dual probability integral transform (PIT) calibration and existing discrete calibration notions, clarifying implication and independence relationships among various calibration types, and providing reproducible methods for generating illustrative examples and counterexamples.

calibrationclassificationhierarchical relations

Fundamentals of Regression

Nov 27, 2025
MA
Miguel A. Mendez
🏛️ von Karman Institute for Fluid Dynamics

This paper addresses the limited physical interpretability and poor generalizability of conventional data-driven regression methods. To this end, we propose a physics-informed regression modeling paradigm that systematically embeds first-principles constraints, differential equation priors, and conservation laws into statistical regression, curve fitting, and supervised learning frameworks—thereby enabling deep integration of machine learning with classical numerical methods (e.g., finite differences, spectral methods). Our key contributions are threefold: (i) we establish a systematic taxonomy tracing the evolution of regression from purely statistical to physics-guided formulations; (ii) we construct a theoretical bridge linking computational science and scientific machine learning; and (iii) the proposed framework significantly enhances model extrapolation capability, robustness, and physical consistency. As a result, it offers an interpretable, verifiable, and cross-disciplinary modeling paradigm for complex scientific and engineering problems.

Combines machine learning and numerical methods for physicsIt integrates physical knowledge with machine learning methodsRegression moves from data-driven to physics-informed formalism

kNN Algorithm for Conditional Mean and Variance Estimation with Automated Uncertainty Quantification and Variable Selection

Feb 02, 2024
MM
Marcos Matabuena
🏛️ Harvard T.H. Chan School of Public Health | University of Santiago de Compostela | University of California, Los Angeles

This paper addresses the challenge of conditional distribution reconstruction under high-dimensional covariates. We propose a novel k-nearest neighbors (kNN) semi-parametric regression method that jointly estimates the conditional mean and variance to reconstruct the conditional density function and quantify predictive uncertainty. Our key contributions are: (1) the first semi-parametric ROC curve estimation framework within kNN; (2) a theoretically grounded, adaptive k-selection algorithm; and (3) a conditional distribution modeling framework integrated with variable selection. Under low-dimensional structural assumptions—such as intrinsic dimensionality or sparsity—we establish consistency and derive the optimal nonparametric convergence rate. Simulation studies demonstrate substantial improvements over conventional kNN methods in estimation accuracy and uncertainty calibration. Furthermore, two real-world biomedical applications confirm the method’s robustness and practical utility in high-dimensional settings.

Automating variance selection to improve empirical performanceEstimating conditional mean and variance via k-NN regressionReconstructing conditional distributions for generative models

Optimizing Estimators of Squared Calibration Errors in Classification

Oct 09, 2024
SG
Sebastian G. Gruber
🏛️ German Cancer Consortium (DKTK) | German Cancer Research Center (DKFZ) | Goethe University Frankfurt | Inria | Ecole Normale Supérieure | PSL Research University

Existing calibration error estimation lacks differentiable, optimizable estimators, hindering end-to-end calibration optimization. Method: We formulate the squared calibration error estimation as a regression task over i.i.d. sample pairs, adopting mean-squared error (MSE) as the risk criterion. Leveraging the bilinear structure of the squared calibration error, we employ kernel ridge regression with joint hyperparameter optimization within a novel train-validation-test estimation pipeline. Contribution/Results: This work establishes the first unified risk-based framework for calibration error estimation; reformulates canonical calibration error estimation as a learnable, differentiable regression problem; and introduces a principled three-stage estimation protocol. Evaluated on standard image classification benchmarks, our estimator achieves significantly higher accuracy than state-of-the-art methods. It is the first practical, end-to-end optimizable estimator for canonical calibration error, enabling gradient-based calibration refinement.

Improving classifier calibration trustworthinessOptimizing squared calibration error estimatorsSelecting and tuning calibration estimators

Estimating and evaluating counterfactual prediction models

Aug 24, 2023
CB
Christopher B. Boyer
🏛️ Cleveland Clinic Research | Case Western Reserve University | Harvard T.H. Chan School of Public Health | Richard A. and Susan F. Smith Center for Outcomes Research | Beth Israel Deaconess Medical Center | Brown University School of Public Health

Counterfactual prediction under evolving intervention policies or hypothetical decision scenarios remains challenging due to unobservable potential outcomes, hindering model identifiability, evaluation, and generalization. Method: We propose the first systematic theoretical framework addressing this challenge—comprising (i) identifiability conditions for counterfactual prediction models, (ii) a performance evaluation system targeting loss, AUC, and calibration, and (iii) robust hyperparameter selection under model misspecification. Our approach integrates causal inference principles, doubly robust estimation, and loss-driven evaluation metric design. Contribution/Results: Validated via simulation studies and a real-world clinical application—cardiovascular risk prediction in statin-naïve populations—the framework significantly improves out-of-distribution generalization and clinical decision reliability in counterfactual settings.

Estimating counterfactual prediction models under different treatment policiesEvaluating model performance without observed potential outcomesProviding valid performance estimates under model misspecification

Latest Papers

What's happening recently
View more

Current evaluation practices for supervised learning models are often misleading due to an overreliance on single aggregate metrics, which neglect the alignment among data characteristics, task objectives, and real-world application contexts. This work reframes model evaluation as a context-dependent, decision-oriented process and systematically investigates—through controlled experiments—the impact of dataset properties, validation strategies, class imbalance, and asymmetric error costs on evaluation outcomes. Leveraging diverse benchmark datasets, multiple validation protocols, and multidimensional performance measures, the study uncovers common pitfalls such as the accuracy paradox, data leakage, and metric misuse. It proposes a structured evaluation framework explicitly aligned with operational goals, offering principled guidance for developing more robust, reliable, and trustworthy supervised learning systems.

class imbalancemodel evaluationperformance metrics

This work proposes a unified inference framework based on parametric programming to address the bias in statistical inference for regression coefficients following Lasso variable selection in generalized linear models. By locally linearizing the maximum likelihood estimator, the method constructs a linear relationship between pseudo-responses and covariates, thereby extending parametric programming—previously limited to Gaussian settings—to non-Gaussian response distributions. This extension enables valid post-selection inference across a range of exponential family models, including logistic and Poisson regression. The approach maintains computational efficiency while substantially improving inferential accuracy. Simulation studies demonstrate that, in non-Gaussian settings, the proposed method effectively corrects the naive inference that ignores the selection process and achieves higher statistical efficiency compared to existing approaches such as the polyhedral method.

generalized linear modelsLassopost-selection inference

meval: A Statistical Toolbox for Fine-Grained Model Performance Analysis

Dec 19, 2025
DS
Dishantkumar Sutariya
🏛️ Fraunhofer Institute for Digital Medicine MEVIS

This study addresses the lack of statistical rigor in evaluating medical AI models across patient- and imaging-based stratifications. We propose the first end-to-end subgroup analysis framework. Methodologically, it integrates Bayesian posterior estimation with bootstrap resampling to quantify metric uncertainty, applies the Benjamini–Hochberg procedure for multiplicity correction, and introduces an adaptive subgroup scanning algorithm to automatically detect statistically significant and clinically relevant performance disparities. The framework explicitly handles challenges including small-sample subgroups, base-rate imbalance, and intersectional subgroup testing. Evaluated on ISIC2020 (skin lesion classification) and MIMIC-CXR (chest radiograph diagnosis), it successfully identifies interpretable fairness-critical subgroups—e.g., demographic or anatomical cohorts exhibiting systematic underperformance—thereby substantially enhancing both the statistical reliability and clinical interpretability of model evaluation.

Addresses statistically rigorous analysis of model performance across subgroups.Identifies interesting subgroups in intersectional analyses among many combinations.Selects appropriate metrics for valid comparisons with varying sample sizes.

This work addresses the issue of miscalibration in probabilistic regression, where predictive distributions often exhibit overconfidence—particularly problematic in safety-critical applications—due to an emphasis on informativeness at the expense of calibration. The authors propose a nonparametric post-hoc calibration method based on conditional kernel mean embeddings, which effectively corrects calibration bias without requiring parametric assumptions about the error structure. By introducing a novel characteristic kernel tailored for real-valued targets, the approach overcomes limitations of existing techniques that rely either on weak calibration criteria or strong parametric assumptions. The method achieves efficient empirical distribution calibration with a computational complexity of $\mathcal{O}(n \log n)$. Extensive experiments across diverse regression benchmarks and models demonstrate that the proposed technique significantly outperforms current post-hoc calibration methods, substantially enhancing the reliability of predictive distributions.

calibrationdistribution regressionnonparametric methods

Multivariate probabilistic forecasting often struggles to balance calibration and informativeness, particularly due to the absence of explicit calibration mechanisms during training. This work proposes a regularization method based on a pre-ranking function that directly optimizes the calibration of multivariate distributional regression models during training. Innovatively, it incorporates principal component analysis (PCA) to construct the pre-ranking strategy, effectively capturing the dominant directions of dependence in the predicted distribution. To the best of our knowledge, this is the first approach to employ a pre-ranking function for calibration-aware regularization in the training phase, and it uniquely reveals misspecifications in dependency structures that existing methods fail to detect. Evaluated across 18 real-world and simulated multi-output datasets, the proposed method significantly improves multivariate calibration while preserving predictive accuracy, demonstrating its effectiveness and practical utility.

distributional regressionmultivariate calibrationpre-rank functions

Hot Scholars

MH

Marco Hutter

Professor of Robotics, ETH Zurich
Legged RoboticsRoboticsControl
CJ

Chenfanfu Jiang

Professor, UCLA
Computer GraphicsComputer VisionEmbodied AIRobotics
KK

Kento Kawaharazuka

The University of Tokyo
HumanoidBiomimeticsTendon-drivenSoft Robotics
NN

Nassir Navab

Professor of Computer Science, Technische Universität München
WW

Wenping Wang

Texas A&M University
Computer GraphicsGeometric Computing