Score
Designs, builds, and analyzes methods, evaluation protocols, and data-processing or training procedures that detect, measure, and address violations of the independent-and-identically-distributed (i.i.d.) assumption in datasets. This includes techniques for identifying distribution shifts and dependencies (e.g., covariate/label shift, temporal or spatial correlation, sample heterogeneity, concept drift), and for mitigating their effects via robust training, reweighting/adaptation, specialized evaluation, or other corrective strategies.
This survey systematically addresses distribution shift between training and deployment in machine learning, focusing on two fundamental challenges: covariate shift (changes in input feature distributions) and concept shift (changes in semantic or class-conditional label distributions). We formalize and unify shift taxonomy, integrating techniques—including distribution shift detection, uncertainty estimation, domain adaptation, anomaly identification, causal inference, and invariant representation learning—within a cohesive framework bridging statistical learning and deep learning. Our key contributions include: (i) a novel robust modeling framework designed to handle heterogeneous shift types; (ii) the first systematic taxonomy covering out-of-distribution (OOD) scenarios; and (iii) a critical analysis revealing limitations of existing methods in jointly mitigating multiple concurrent shifts and generalizing to unseen classes. We establish principled evaluation criteria and outline future research directions—particularly addressing compound shifts and semantic evolution—thereby filling a critical gap in prior surveys, which largely overlook real-world deployment complexities involving intertwined and dynamically evolving shifts.
This study investigates the impact of expanding the historical training window on predictive performance and algorithmic fairness under temporal distribution shift—including covariate and concept drift. Using simulation experiments and empirical student retention prediction across multi-institution, multi-year educational datasets, we find that the common assumption “more data is better” fails under concept-drift-dominant regimes: extending the training window degrades overall accuracy and exacerbates fairness disparities across sociodemographic groups—particularly when marginalized subpopulations experience heterogeneous concept drift, leading to nonlinear bias amplification. Our key contributions are: (i) identifying concept drift—not covariate drift—as the primary driver of performance degradation; and (ii) the first systematic characterization of its nonlinear, fairness-amplifying mechanism. These findings provide both theoretical grounding and practical guidance for model update strategies in dynamic, real-world deployment environments.
本文探讨了在数据不完美条件下的机器学习挑战,并通过信息损失、经验风险偏差等机制组织代表性方法,如重建生成、再平衡与表示校准等。
Distribution shift in tabular data lacks empirical grounding, and existing robust learning methods—such as distributionally robust optimization (DRO)—frequently fail in real-world settings. Prior work predominantly assumes covariate shift (X-shift), whereas empirical analysis reveals that label-conditional shift (Y|X-shift) is both more prevalent and dominant in practice. Method: The authors construct a comprehensive benchmark platform spanning five real-world tabular datasets and 60,000 experimental configurations, enabling large-scale ablation studies on robustness. Contribution/Results: (1) Theory-driven methods like DRO offer no consistent advantage over empirical risk minimization (ERM); (2) implementation details—including model selection, hyperparameter tuning, and imbalance handling—exert far greater influence on robustness than the design of ambiguity sets; (3) a data-driven, inductive paradigm should supplant prior structural assumptions. These findings challenge the conventional paradigm in distributionally robust learning and provide reproducible, empirically grounded guidance for co-optimizing algorithms and data.
This study addresses the bottleneck of concept drift detection in large-scale e-commerce data characterized by hundreds of millions of rows and high-dimensional features. Leveraging Apache Spark, it evaluates multi-column two-sample drift detection methods, with particular emphasis on the scalability of a distributed Maximum Mean Discrepancy (MMD) algorithm integrated with Random Fourier Features. A synthetic injection benchmark comprising 137.5 million rows is constructed, revealing a critical limitation wherein the Kolmogorov-Smirnov test fails due to statistical saturation in identifier-like columns. Experimental results demonstrate that under strong drift conditions, the proposed method achieves a Pearson correlation coefficient of 0.940, a true positive rate of 80.4%, and a false positive rate of merely 3.2%. However, its sensitivity remains constrained in weak drift scenarios.
In AI deployment for medical imaging, data distribution shifts frequently cause abrupt performance degradation and increased misdiagnosis risk. Existing methods can only detect the presence of shift but fail to identify its specific type—e.g., covariate shift, prior (concept) shift, or compound shift—hindering root-cause analysis and targeted mitigation. This paper proposes the first unsupervised framework for data shift type identification. It introduces a novel joint shift detection mechanism that synergistically leverages self-supervised encoder representations and task-model outputs. By integrating feature distribution comparison, unsupervised clustering, and multimodal image modeling, the method achieves high-accuracy shift-type discrimination across three major imaging modalities—chest X-ray, mammography, and fundus photography—and five realistic shift scenarios. Evaluated on four large public medical imaging datasets, it significantly enhances the robustness and interpretability of clinical AI systems.
When predictive models are deployed in new environments, their performance often degrades due to covariate shift and selective labeling, which jointly obscure accurate assessment of the true target risk. This work proposes an unbiased risk estimation method that integrates double machine learning with influence functions to simultaneously address both sources of bias for the first time. The approach is model-agnostic and compatible with general loss functions, constructing a corrected target risk estimator via nonparametric and plug-in techniques. Experiments on eICU electronic health record data demonstrate that the proposed method significantly outperforms baselines that handle only one type of bias or naively combine existing approaches, yielding more accurate tracking of the true target risk.
This study addresses a critical limitation in existing design-based simulations used to evaluate inference methods, which often overstate bias induced by spatial correlation due to unrealistic data-generating mechanisms. In particular, share-shift designs that fix outcomes and resample shocks conflate true treatment effects with error dependence structures, leading to misleading assessments. To remedy this, the paper proposes an improved simulation framework that more accurately models error dependence and avoids spurious entanglement between treatment effects and error terms, thereby better approximating real-world data-generating processes. Integrating resampling techniques with share-shift analysis, the proposed approach substantially enhances the reliability of inference evaluation across multiple empirical applications, underscoring the essential role of aligning simulation designs with genuine underlying mechanisms for valid inference assessment.
This study addresses the performance degradation of machine learning classifiers in dynamic environments caused by concept drift, a phenomenon inadequately captured by conventional evaluation methods that overlook causal dependencies in data, leading to distorted assessments. To overcome this limitation, the authors propose a digital twin framework grounded in Structural Causal Models (SCMs), which, for the first time, leverages SCMs to simulate realistic causal drift. By applying parametric causal interventions, the framework stress-tests classifiers while preserving the underlying structure of the data-generating mechanism. This approach transcends the constraints of traditional statistical or correlation-based evaluations. Experimental results on the OSMH dataset demonstrate that the method effectively uncovers classifier vulnerabilities that remain undetected by standard monitoring techniques.
This study addresses the scarcity of tabular data under covariate shift and the instability of conventional augmentation methods caused by fitting outdated source distributions. To this end, we propose the IGDPR framework, which pioneers the integration of invariant potential functions into the diffusion sampling process to align stable decision boundaries and generate task-relevant samples. Furthermore, prototype clustering and reweighting strategies are incorporated to filter noise and assess sample reliability, thereby mitigating overfitting. Experimental results demonstrate that the proposed framework significantly improves synthetic data quality in real-world scenarios, effectively enhancing model robustness and generalization to unseen environments.
本文解决了机器学习中分布偏移下的泛化问题,通过引入γ*-概念偏移和熵最优传输方法,提出了统一的误差界,并开发了能够估计这些偏移的算法。