target risk estimation

Designs and evaluates methods to estimate a model’s expected predictive loss (risk) in the intended target environment, including empirical and statistical estimators used for pre-deployment risk evaluation and off-distribution assessment. Builds estimators and analyses that correct for covariate distribution shift, selective label observability, and arbitrary loss functions, and that quantify bias and uncertainty in the risk estimate.

targetriskestimation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.39
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

When predictive models are deployed in new environments, their performance often degrades due to covariate shift and selective labeling, which jointly obscure accurate assessment of the true target risk. This work proposes an unbiased risk estimation method that integrates double machine learning with influence functions to simultaneously address both sources of bias for the first time. The approach is model-agnostic and compatible with general loss functions, constructing a corrected target risk estimator via nonparametric and plug-in techniques. Experiments on eICU electronic health record data demonstrate that the proposed method significantly outperforms baselines that handle only one type of bias or naively combine existing approaches, yielding more accurate tracking of the true target risk.

covariate shiftdistribution shiftmodel evaluation

Estimating and evaluating counterfactual prediction models

Aug 24, 2023
CB
Christopher B. Boyer
🏛️ Cleveland Clinic Research | Case Western Reserve University | Harvard T.H. Chan School of Public Health | Richard A. and Susan F. Smith Center for Outcomes Research | Beth Israel Deaconess Medical Center | Brown University School of Public Health

Counterfactual prediction under evolving intervention policies or hypothetical decision scenarios remains challenging due to unobservable potential outcomes, hindering model identifiability, evaluation, and generalization. Method: We propose the first systematic theoretical framework addressing this challenge—comprising (i) identifiability conditions for counterfactual prediction models, (ii) a performance evaluation system targeting loss, AUC, and calibration, and (iii) robust hyperparameter selection under model misspecification. Our approach integrates causal inference principles, doubly robust estimation, and loss-driven evaluation metric design. Contribution/Results: Validated via simulation studies and a real-world clinical application—cardiovascular risk prediction in statin-naïve populations—the framework significantly improves out-of-distribution generalization and clinical decision reliability in counterfactual settings.

Estimating counterfactual prediction models under different treatment policiesEvaluating model performance without observed potential outcomesProviding valid performance estimates under model misspecification

Generalization and Informativeness of Weighted Conformal Risk Control Under Covariate Shift

Jan 20, 2025
MZ
Matteo Zecchin
🏛️ King's College London | University College London | Technion—Israel Institute of Technology

This paper addresses the inefficiency (i.e., prediction set size) and generalization of weighted conformal risk control (W-CRC) under covariate shift. Methodologically, it integrates importance weighting, conformal prediction, and statistical learning theory to derive a computable upper bound on inefficiency during training—explicitly linking it to the base predictor’s generalization error, the degree of distribution shift, and sample size. The key contribution is the first quantitative characterization of how W-CRC efficiency depends on the base predictor’s generalization capability, and the theoretical revelation that prediction set informativeness decays with increasing shift severity. Empirical validation on fingerprint-based indoor localization demonstrates that the derived bound accurately captures the trend of prediction set size under varying shift levels, confirming both theoretical soundness and practical utility.

Covariate ShiftModel GeneralizationWeighted Consistency Risk Control

This work addresses the challenge of predicting performance changes when a source-domain model is replaced by a new one. To this end, the authors propose TRACE, a novel framework that, for the first time, decomposes the risk difference between two models under covariate shift into four interpretable components: two generalization gaps, a model change penalty, and a covariate shift penalty. The framework establishes a computable upper bound to diagnose the causes of performance degradation. TRACE estimates model sensitivity via high-quantile input gradients, quantifies data distribution shift using either optimal transport (OT) or maximum mean discrepancy (MMD), and measures model change through output distances on target samples. Experiments demonstrate that TRACE’s diagnostic scores exhibit strong monotonic correlation with actual performance degradation and achieve superior performance in deployment gating, as measured by AUROC and AUPRC, thereby enabling label-efficient and safe model replacement.

covariate shiftdistribution shiftmodel replacement

On the Variance, Admissibility, and Stability of Empirical Risk Minimization

May 29, 2023
GK
Gil Kur
🏛️ MIT | Tel Aviv University

This paper identifies an inherent suboptimality of Empirical Risk Minimization (ERM) under squared loss: its bias term dominates the estimation error, preventing attainment of the minimax optimal convergence rate; in contrast, the variance term achieves the minimax rate—a fact previously unverified rigorously. To address this, the authors establish, for the first time under random design, a non-asymptotic, sharp upper bound proving the minimax optimality of ERM’s variance term. They unify and extend Chatterjee’s admissibility theorem and the Caponnetto–Rakhlin stability result to realistic random-design settings. Furthermore, they systematically characterize the intrinsic irregularity of the empirical loss landscape for non-Donsker function classes. Integrating bias–variance decomposition, empirical process theory, and probabilistic analysis, the work delivers the most comprehensive and rigorous non-asymptotic risk decomposition for ERM to date, substantially advancing its theoretical foundations.

ERM cannot be ruled out as an optimal estimation method.ERM exhibits irregular loss landscape in non-Donsker regimes.ERM's suboptimality stems from large bias, not variance error.

Latest Papers

What's happening recently
View more

Traditional cross-validation in spatial prediction suffers from biased risk estimation due to distributional mismatches between validation and deployment tasks, including covariate shift and task difficulty shift. This work proposes Target-Weighted Cross-Validation (TWCV), which for the first time incorporates task distribution alignment into spatial prediction by calibrating weights to match the target-domain distribution and enhancing task difficulty diversity through spatial buffering-based resampling. By integrating importance-weighted risk estimation with task descriptor modeling, TWCV substantially reduces estimation bias. Empirical evaluations on both synthetic data and real-world environmental pollution mapping demonstrate that the method yields more accurate and nearly unbiased estimates of deployment risk compared to existing spatial cross-validation approaches.

covariate shiftcross-validationdistribution shift

This work addresses the challenge of risk prediction when acquiring true outcomes is prohibitively expensive, allowing labels for only a subset of samples. The authors propose a surrogate-assisted optimal sampling framework that, under a fixed annotation budget, leverages covariates, surrogate variables, and an initial estimator to construct a sampling strategy minimizing the expected out-of-sample cross-entropy loss. Coupled with an inverse probability weighted cross-entropy estimator for model training, this approach achieves—without requiring access to true responses during design—theoretically optimal sampling. It uniquely guarantees predictive optimality, robustness to surrogate misspecification, and stability in settings with rare outcomes. Both theoretical analysis and empirical experiments demonstrate that the method significantly outperforms existing approaches, particularly when the surrogate is imperfect or the event of interest is rare.

budget allocationmeasurement constraintsoptimal sampling

Hot Scholars

MS

Masashi Sugiyama

Director, RIKEN Center for Advanced Intelligence Project / Professor, The University of Tokyo
Machine LearningData MiningArtificial Intelligence
XZ

Xingyu Zhao

Associate Professor, University of Warwick
Software ReliabilitySafe AIBayesian InferenceProbabilistic Model Checking
MI

Michael I. Jordan

Professor of Electrical Engineering and Computer Sciences and Professor of Statistics, UC Berkeley
machine learningcomputer sciencestatisticsartificial intelligence
AM

Andrea Montanari

John D. and Sigrid Banks Professor, Statistics and Mathematics, Stanford University
statisticsmachine learningprobability theoryinformation theory
DD

Darinka Dentcheva

Professor of Mathematics, Stevens Institute of Technology
nonlinear optimizationstochastic optimizationriskconvex analysis