Score
Designs and implements methods to detect, measure, and quantify changes in the distribution of input covariates between datasets or over time, and to evaluate how those changes affect model outputs and performance. Builds estimators and statistical tests for covariate shift, derives error bounds (e.g., for projection or importance-weighted risk), and produces assessments used to select, weight, or diagnose auxiliary data to reduce shift-induced errors.
This survey systematically addresses distribution shift between training and deployment in machine learning, focusing on two fundamental challenges: covariate shift (changes in input feature distributions) and concept shift (changes in semantic or class-conditional label distributions). We formalize and unify shift taxonomy, integrating techniques—including distribution shift detection, uncertainty estimation, domain adaptation, anomaly identification, causal inference, and invariant representation learning—within a cohesive framework bridging statistical learning and deep learning. Our key contributions include: (i) a novel robust modeling framework designed to handle heterogeneous shift types; (ii) the first systematic taxonomy covering out-of-distribution (OOD) scenarios; and (iii) a critical analysis revealing limitations of existing methods in jointly mitigating multiple concurrent shifts and generalizing to unseen classes. We establish principled evaluation criteria and outline future research directions—particularly addressing compound shifts and semantic evolution—thereby filling a critical gap in prior surveys, which largely overlook real-world deployment complexities involving intertwined and dynamically evolving shifts.
When predictive models are deployed in new environments, their performance often degrades due to covariate shift and selective labeling, which jointly obscure accurate assessment of the true target risk. This work proposes an unbiased risk estimation method that integrates double machine learning with influence functions to simultaneously address both sources of bias for the first time. The approach is model-agnostic and compatible with general loss functions, constructing a corrected target risk estimator via nonparametric and plug-in techniques. Experiments on eICU electronic health record data demonstrate that the proposed method significantly outperforms baselines that handle only one type of bias or naively combine existing approaches, yielding more accurate tracking of the true target risk.
This work addresses the challenge of predicting performance changes when a source-domain model is replaced by a new one. To this end, the authors propose TRACE, a novel framework that, for the first time, decomposes the risk difference between two models under covariate shift into four interpretable components: two generalization gaps, a model change penalty, and a covariate shift penalty. The framework establishes a computable upper bound to diagnose the causes of performance degradation. TRACE estimates model sensitivity via high-quantile input gradients, quantifies data distribution shift using either optimal transport (OT) or maximum mean discrepancy (MMD), and measures model change through output distances on target samples. Experiments demonstrate that TRACE’s diagnostic scores exhibit strong monotonic correlation with actual performance degradation and achieve superior performance in deployment gating, as measured by AUROC and AUPRC, thereby enabling label-efficient and safe model replacement.
To address the challenge of unsupervised model performance estimation under covariate shift—where ground-truth labels are unavailable or delayed post-deployment—this paper proposes the Probability-Adaptive Performance Estimation (PAPE) framework. PAPE requires neither access to true labels nor knowledge of the original model’s architecture or feature representations; it operates solely on the model’s probabilistic outputs and confidence scores. By jointly leveraging density ratio estimation and performance generalization bound theory, PAPE models prediction distributions and applies adaptive reweighting to yield unbiased estimates of arbitrary classification metrics—without assuming a specific shift form or resorting to feature learning or generative modeling. Extensive evaluation across 900+ real-world census dataset–model combinations demonstrates that PAPE reduces mean absolute error by 37% compared to state-of-the-art proxy metrics and drift detection methods, significantly enhancing the reliability and generality of model monitoring in production environments.
In AI deployment for medical imaging, data distribution shifts frequently cause abrupt performance degradation and increased misdiagnosis risk. Existing methods can only detect the presence of shift but fail to identify its specific type—e.g., covariate shift, prior (concept) shift, or compound shift—hindering root-cause analysis and targeted mitigation. This paper proposes the first unsupervised framework for data shift type identification. It introduces a novel joint shift detection mechanism that synergistically leverages self-supervised encoder representations and task-model outputs. By integrating feature distribution comparison, unsupervised clustering, and multimodal image modeling, the method achieves high-accuracy shift-type discrimination across three major imaging modalities—chest X-ray, mammography, and fundus photography—and five realistic shift scenarios. Evaluated on four large public medical imaging datasets, it significantly enhances the robustness and interpretability of clinical AI systems.
This paper addresses the unreliability of causal and predictive parameter estimation under covariate shift. We propose a fully automated debiasing machine learning framework that eliminates regularization bias solely through parameter definition—without requiring explicit bias modeling. Our approach innovatively integrates training and target data within a unified debiasing mechanism, combining data fusion, high-dimensional statistical inference, doubly robust estimation, and the difference-in-differences (DID) principle—all under an unconfoundedness assumption. We establish theoretical guarantees of consistency and asymptotic normality. In simulation studies and an empirical analysis of minimum wage effects on teenage employment, our method reduces estimation bias by over 40% on average compared to benchmark approaches, while substantially improving estimation accuracy and robustness.
This study addresses a critical limitation in existing design-based simulations used to evaluate inference methods, which often overstate bias induced by spatial correlation due to unrealistic data-generating mechanisms. In particular, share-shift designs that fix outcomes and resample shocks conflate true treatment effects with error dependence structures, leading to misleading assessments. To remedy this, the paper proposes an improved simulation framework that more accurately models error dependence and avoids spurious entanglement between treatment effects and error terms, thereby better approximating real-world data-generating processes. Integrating resampling techniques with share-shift analysis, the proposed approach substantially enhances the reliability of inference evaluation across multiple empirical applications, underscoring the essential role of aligning simulation designs with genuine underlying mechanisms for valid inference assessment.
This study addresses the pervasive issue of measurement error in both outcome variables and multiple covariates within routinely collected biomedical data, such as electronic health records, which, if uncorrected, can induce analytical bias and misinform clinical decisions. For the first time within a tutorial framework, it systematically reviews and empirically compares several methods capable of simultaneously correcting measurement error in both outcomes and multiple covariates—including regression calibration, SIMEX, instrumental variable approaches, and modeling strategies leveraging validation subsamples. Through a unified illustrative example and publicly available code, the work not only clarifies the relative performance of these methods in real-world data to guide researchers’ methodological choices but also establishes a reproducible end-to-end analytical pipeline and highlights promising directions for future research.