Score
Designs and evaluates methods to estimate a model’s expected predictive loss (risk) in the intended target environment, including empirical and statistical estimators used for pre-deployment risk evaluation and off-distribution assessment. Builds estimators and analyses that correct for covariate distribution shift, selective label observability, and arbitrary loss functions, and that quantify bias and uncertainty in the risk estimate.
When predictive models are deployed in new environments, their performance often degrades due to covariate shift and selective labeling, which jointly obscure accurate assessment of the true target risk. This work proposes an unbiased risk estimation method that integrates double machine learning with influence functions to simultaneously address both sources of bias for the first time. The approach is model-agnostic and compatible with general loss functions, constructing a corrected target risk estimator via nonparametric and plug-in techniques. Experiments on eICU electronic health record data demonstrate that the proposed method significantly outperforms baselines that handle only one type of bias or naively combine existing approaches, yielding more accurate tracking of the true target risk.
Counterfactual prediction under evolving intervention policies or hypothetical decision scenarios remains challenging due to unobservable potential outcomes, hindering model identifiability, evaluation, and generalization. Method: We propose the first systematic theoretical framework addressing this challenge—comprising (i) identifiability conditions for counterfactual prediction models, (ii) a performance evaluation system targeting loss, AUC, and calibration, and (iii) robust hyperparameter selection under model misspecification. Our approach integrates causal inference principles, doubly robust estimation, and loss-driven evaluation metric design. Contribution/Results: Validated via simulation studies and a real-world clinical application—cardiovascular risk prediction in statin-naïve populations—the framework significantly improves out-of-distribution generalization and clinical decision reliability in counterfactual settings.
This paper addresses the inefficiency (i.e., prediction set size) and generalization of weighted conformal risk control (W-CRC) under covariate shift. Methodologically, it integrates importance weighting, conformal prediction, and statistical learning theory to derive a computable upper bound on inefficiency during training—explicitly linking it to the base predictor’s generalization error, the degree of distribution shift, and sample size. The key contribution is the first quantitative characterization of how W-CRC efficiency depends on the base predictor’s generalization capability, and the theoretical revelation that prediction set informativeness decays with increasing shift severity. Empirical validation on fingerprint-based indoor localization demonstrates that the derived bound accurately captures the trend of prediction set size under varying shift levels, confirming both theoretical soundness and practical utility.
This work addresses the challenge of predicting performance changes when a source-domain model is replaced by a new one. To this end, the authors propose TRACE, a novel framework that, for the first time, decomposes the risk difference between two models under covariate shift into four interpretable components: two generalization gaps, a model change penalty, and a covariate shift penalty. The framework establishes a computable upper bound to diagnose the causes of performance degradation. TRACE estimates model sensitivity via high-quantile input gradients, quantifies data distribution shift using either optimal transport (OT) or maximum mean discrepancy (MMD), and measures model change through output distances on target samples. Experiments demonstrate that TRACE’s diagnostic scores exhibit strong monotonic correlation with actual performance degradation and achieve superior performance in deployment gating, as measured by AUROC and AUPRC, thereby enabling label-efficient and safe model replacement.
This paper identifies an inherent suboptimality of Empirical Risk Minimization (ERM) under squared loss: its bias term dominates the estimation error, preventing attainment of the minimax optimal convergence rate; in contrast, the variance term achieves the minimax rate—a fact previously unverified rigorously. To address this, the authors establish, for the first time under random design, a non-asymptotic, sharp upper bound proving the minimax optimality of ERM’s variance term. They unify and extend Chatterjee’s admissibility theorem and the Caponnetto–Rakhlin stability result to realistic random-design settings. Furthermore, they systematically characterize the intrinsic irregularity of the empirical loss landscape for non-Donsker function classes. Integrating bias–variance decomposition, empirical process theory, and probabilistic analysis, the work delivers the most comprehensive and rigorous non-asymptotic risk decomposition for ERM to date, substantially advancing its theoretical foundations.
Traditional cross-validation in spatial prediction suffers from biased risk estimation due to distributional mismatches between validation and deployment tasks, including covariate shift and task difficulty shift. This work proposes Target-Weighted Cross-Validation (TWCV), which for the first time incorporates task distribution alignment into spatial prediction by calibrating weights to match the target-domain distribution and enhancing task difficulty diversity through spatial buffering-based resampling. By integrating importance-weighted risk estimation with task descriptor modeling, TWCV substantially reduces estimation bias. Empirical evaluations on both synthetic data and real-world environmental pollution mapping demonstrate that the method yields more accurate and nearly unbiased estimates of deployment risk compared to existing spatial cross-validation approaches.
This work addresses the challenge of risk prediction when acquiring true outcomes is prohibitively expensive, allowing labels for only a subset of samples. The authors propose a surrogate-assisted optimal sampling framework that, under a fixed annotation budget, leverages covariates, surrogate variables, and an initial estimator to construct a sampling strategy minimizing the expected out-of-sample cross-entropy loss. Coupled with an inverse probability weighted cross-entropy estimator for model training, this approach achieves—without requiring access to true responses during design—theoretically optimal sampling. It uniquely guarantees predictive optimality, robustness to surrogate misspecification, and stability in settings with rare outcomes. Both theoretical analysis and empirical experiments demonstrate that the method significantly outperforms existing approaches, particularly when the surrogate is imperfect or the event of interest is rare.