Score
Designs and implements methods, tools, and evaluations that simulate, detect, and reduce model performance degradation caused by differences between training and deployment data distributions. This includes building realistic distribution-shift simulators, data augmentations and in-distribution counterfactual generators, and shift-specific robustness tests and mitigation strategies.
This survey systematically addresses distribution shift between training and deployment in machine learning, focusing on two fundamental challenges: covariate shift (changes in input feature distributions) and concept shift (changes in semantic or class-conditional label distributions). We formalize and unify shift taxonomy, integrating techniques—including distribution shift detection, uncertainty estimation, domain adaptation, anomaly identification, causal inference, and invariant representation learning—within a cohesive framework bridging statistical learning and deep learning. Our key contributions include: (i) a novel robust modeling framework designed to handle heterogeneous shift types; (ii) the first systematic taxonomy covering out-of-distribution (OOD) scenarios; and (iii) a critical analysis revealing limitations of existing methods in jointly mitigating multiple concurrent shifts and generalizing to unseen classes. We establish principled evaluation criteria and outline future research directions—particularly addressing compound shifts and semantic evolution—thereby filling a critical gap in prior surveys, which largely overlook real-world deployment complexities involving intertwined and dynamically evolving shifts.
This work addresses the severe performance degradation and erroneous detection of aerial vision models under weather-induced distribution shifts. To this end, we introduce MDS-A—the first multi-distribution-shift benchmark for aerial imagery—featuring high-fidelity synthetic training data generated in Unreal Engine under six controlled meteorological conditions, a mixed-weather test set, and comprehensive annotations. We propose EDR (Error Detection and Recovery), a knowledge-driven framework enabling test-time uncertainty modeling and lightweight self-adaptation. MDS-A is the first benchmark to support fine-grained out-of-distribution (OOD) attribution analysis and standardized evaluation across multidimensional weather shifts. Experiments show that mainstream YOLOv5/v8 models suffer 32–68% mAP drops across weather domains; with EDR, erroneous detection accuracy reaches 89.7%, and online model adaptation is effectively triggered.
This paper investigates the robustness degradation of machine learning models under concurrent distribution shifts—specifically, the co-occurrence of domain shift and spurious correlations. To this end, we establish a comprehensive benchmark spanning eight datasets, 168 source–target domain pairs, and 26 algorithms, involving over 100,000 model training and evaluation runs. We propose a multi-source–multi-target shift construction framework and a statistical attribution analysis methodology. Our large-scale empirical study is the first to systematically quantify the compounding effect of concurrent shifts; reveals positive cross-shift generalization transferability; and demonstrates that heuristic data augmentation consistently outperforms large-model zero-shot inference—achieving state-of-the-art average robustness on both synthetic and real-world benchmarks. Crucially, we identify a consistent cross-shift pattern in generalization improvement, providing both theoretical grounding and practical guidance for robust modeling in complex, realistic deployment scenarios.
This paper addresses the reliable detection of post-deployment performance degradation (PDD) in unlabeled model-serving scenarios. We formally define the PDD monitoring task as distinguishing benign distributional shifts from genuine performance deterioration. To this end, we propose D3M—a label-free, gradient-free monitoring framework that leverages predictive disagreement across multiple models. We theoretically establish its low false-positive rate under non-degrading shifts and provide sample-complexity guarantees. By unifying theoretical analysis with empirical risk estimation, D3M achieves significant improvements over state-of-the-art baselines on standard benchmarks and a large-scale real-world internal medicine dataset. Our method delivers a verifiable, automated alerting mechanism for performance degradation in high-stakes machine learning systems.
This work addresses the pervasive distribution shift problem in machine learning–enhanced hybrid simulation. We first establish a formal mathematical modeling and theoretical analysis framework, revealing the root causes of distribution shift and its error-amplification mechanism over long-term simulation. To mitigate shift propagation, we propose the Tangent Space Regularized Estimator (TSRE), which explicitly enforces consistency of the underlying manifold’s tangent space during surrogate model training. We provide rigorous theoretical guarantees showing that TSRE significantly tightens the long-horizon simulation error bound. Extensive experiments on strongly nonlinear reaction–diffusion systems and high-Reynolds-number Navier–Stokes simulations demonstrate that TSRE reduces average prediction error by 42% beyond 100 time steps compared to baseline methods, with especially pronounced gains under severe distribution shift. This work delivers the first theoretically grounded, distributionally robust solution for ML-enhanced simulation.
Distribution shift in tabular data lacks empirical grounding, and existing robust learning methods—such as distributionally robust optimization (DRO)—frequently fail in real-world settings. Prior work predominantly assumes covariate shift (X-shift), whereas empirical analysis reveals that label-conditional shift (Y|X-shift) is both more prevalent and dominant in practice. Method: The authors construct a comprehensive benchmark platform spanning five real-world tabular datasets and 60,000 experimental configurations, enabling large-scale ablation studies on robustness. Contribution/Results: (1) Theory-driven methods like DRO offer no consistent advantage over empirical risk minimization (ERM); (2) implementation details—including model selection, hyperparameter tuning, and imbalance handling—exert far greater influence on robustness than the design of ambiguity sets; (3) a data-driven, inductive paradigm should supplant prior structural assumptions. These findings challenge the conventional paradigm in distributionally robust learning and provide reproducible, empirically grounded guidance for co-optimizing algorithms and data.
This work addresses the limitation of existing static tabular datasets, which lack temporal structure and thus hinder the evaluation of model adaptability under controlled distribution shifts. To overcome this, the authors propose a clustering-based framework that transforms static data into controllable, evolving data streams through cluster-based partitioning and structured perturbations. Integrating the ADWIN drift detector with a sliding-window retraining mechanism, the framework systematically evaluates adaptation strategies across six model families, including tree ensembles and online learners. Experiments on five benchmark datasets for classification and regression demonstrate that the proposed methods—particularly Clustered Local ADWIN—accurately model and efficiently respond to localized drifts in feature space, significantly outperforming baseline approaches.
This work addresses the challenging problem of model performance estimation under distribution shift when ground-truth labels are unavailable. The authors propose Fused Reference Alignment Prediction (FRAP), a novel approach that synergistically combines the generalization capability of external foundation models with the domain-specific expertise of the target task model. FRAP calibrates prediction distributions via temperature scaling, aligns them by minimizing KL divergence, and generates a robust, domain-adaptive pseudo-label reference distribution through confidence-weighted fusion. Extensive experiments demonstrate that FRAP consistently and substantially outperforms existing performance estimation methods across diverse datasets and model architectures, achieving stable and significant improvements.
When predictive models are deployed in new environments, their performance often degrades due to covariate shift and selective labeling, which jointly obscure accurate assessment of the true target risk. This work proposes an unbiased risk estimation method that integrates double machine learning with influence functions to simultaneously address both sources of bias for the first time. The approach is model-agnostic and compatible with general loss functions, constructing a corrected target risk estimator via nonparametric and plug-in techniques. Experiments on eICU electronic health record data demonstrate that the proposed method significantly outperforms baselines that handle only one type of bias or naively combine existing approaches, yielding more accurate tracking of the true target risk.
Current pre-deployment safety evaluations often fail to accurately predict the frequency of undesirable behaviors in large language models during real-world deployment due to insufficient coverage, unrepresentative samples, and susceptibility to being recognized by models as test inputs. This work proposes a deployment simulation method grounded in authentic dialogue prefixes: by fixing historical context and prompting candidate models to generate subsequent responses, it enables auditing of novel alignment failures and estimation of risk incidence rates. The approach leverages publicly available chat data to construct evaluation scenarios, allowing external researchers to conduct realistic safety assessments without access to proprietary logs. Prospective and retrospective experiments on the GPT-5 model series demonstrate that this method significantly outperforms baselines based on adversarial production data, yielding predictions that align more closely with observed misbehavior rates in actual deployment and proving feasible even in complex tool-use settings.