Score
Designs, builds, and evaluates algorithms, tests, and monitoring systems that detect when the input or feature distribution seen by a model or system deviates from an expected or training distribution (including out‑of‑distribution examples and change‑point detection). Develops and applies statistical tests, scoring functions, evaluation metrics, severity/uncertainty estimation, root‑cause analysis, and operational handling policies (alerts, thresholds, retraining or adaptation triggers) to analyze, validate, and respond to observed distribution shifts.
This survey systematically addresses distribution shift between training and deployment in machine learning, focusing on two fundamental challenges: covariate shift (changes in input feature distributions) and concept shift (changes in semantic or class-conditional label distributions). We formalize and unify shift taxonomy, integrating techniques—including distribution shift detection, uncertainty estimation, domain adaptation, anomaly identification, causal inference, and invariant representation learning—within a cohesive framework bridging statistical learning and deep learning. Our key contributions include: (i) a novel robust modeling framework designed to handle heterogeneous shift types; (ii) the first systematic taxonomy covering out-of-distribution (OOD) scenarios; and (iii) a critical analysis revealing limitations of existing methods in jointly mitigating multiple concurrent shifts and generalizing to unseen classes. We establish principled evaluation criteria and outline future research directions—particularly addressing compound shifts and semantic evolution—thereby filling a critical gap in prior surveys, which largely overlook real-world deployment complexities involving intertwined and dynamically evolving shifts.
Real-world machine learning systems operate under dynamic data distribution shifts, rendering conventional risk control methods—predicated on static distributional assumptions—ineffective and incapable of online monitoring for decision risk violations. To address this, we propose the first sequential testing framework grounded in the “betting” paradigm, which strictly controls the false alarm rate (≤ α) under arbitrary, unknown distributional shifts—without assuming prior knowledge of drift type or underlying distributions. Our approach unifies betting-based hypothesis testing, risk-bound modeling, and online streaming statistical inference to enable real-time, robust monitoring of model risk. Extensive experiments demonstrate that the method achieves high sensitivity in detecting risk violations across diverse drift scenarios, while simultaneously delivering rigorous statistical guarantees in both anomaly detection and conformal prediction tasks.
To address model performance degradation caused by data distribution drift and the reliance of existing MLOps retraining pipelines on manual intervention, this paper proposes an automated, adaptive neural network retraining framework. Methodologically, it introduces a novel multi-criteria joint drift detection mechanism—integrating statistical metrics including the Kolmogorov–Smirnov test, Population Stability Index (PSI), and Classifier-Driven (CD) drift detection—combined with online monitoring and lightweight scheduling to dynamically trigger end-to-end retraining upon significant drift. The framework is implemented using a cloud-native architecture for scalable and efficient deployment. Evaluated on multiple benchmark datasets, the proposed approach improves classification accuracy by 12.3%–18.7%, reduces inference latency by 41%, and cuts computational resource consumption by 53%, compared to conventional periodic or single-threshold retraining strategies. These gains significantly enhance model freshness and operational cost-efficiency.
This work addresses the severe performance degradation and erroneous detection of aerial vision models under weather-induced distribution shifts. To this end, we introduce MDS-A—the first multi-distribution-shift benchmark for aerial imagery—featuring high-fidelity synthetic training data generated in Unreal Engine under six controlled meteorological conditions, a mixed-weather test set, and comprehensive annotations. We propose EDR (Error Detection and Recovery), a knowledge-driven framework enabling test-time uncertainty modeling and lightweight self-adaptation. MDS-A is the first benchmark to support fine-grained out-of-distribution (OOD) attribution analysis and standardized evaluation across multidimensional weather shifts. Experiments show that mainstream YOLOv5/v8 models suffer 32–68% mAP drops across weather domains; with EDR, erroneous detection accuracy reaches 89.7%, and online model adaptation is effectively triggered.
This work addresses the limitation of existing safety-critical systems, which typically evaluate only predictive accuracy while lacking rigorous validation of the overall calibration of predicted probability distributions. To bridge this gap, the authors propose a modular calibration testing framework that decouples the calibration process into four interchangeable components: data model, scoring rule, hypothesis formulation, and statistical test procedure. Built upon formal statistical hypothesis testing, the framework provides a single accept/reject decision for the entire predictive distribution. Crucially, it rejects only overly confident predictions while tolerating reasonable deviations, thereby balancing practicality with flexibility. Empirical evaluations on weather forecasting and robotic pose estimation tasks demonstrate that the framework effectively supports reliable deployment in safety-critical applications.
In industrial quality inspection, anomaly detection suffers from poor robustness due to high noise levels and sparse defective samples. To address this, we propose Iterative Refinement of Pseudo-labels (IRP), a self-supervised method that alternately evaluates sample credibility and removes misleading instances under feature-space consistency constraints—effectively purifying the training set dynamically without human annotations and generating high-fidelity self-supervised signals. IRP introduces the novel paradigm of “iterative data refinement,” significantly enhancing model robustness against label noise and cross-domain generalization capability. Evaluated on KSDD2 and MVTec AD benchmarks, IRP consistently outperforms existing unsupervised and self-supervised methods. Notably, under high-noise conditions, it achieves substantial improvements in detection accuracy and reduces false positive rates by over 25%.
This work addresses the challenge of accurately attributing detected change points in multivariate time series to specific subsets of variables. The authors propose a post-hoc, nonparametric testing framework that, after an offline change point has been identified, determines whether the change occurs in one of two pre-specified coordinate blocks or in both. Built upon two-sample nonparametric hypothesis testing, the method offers rigorous theoretical guarantees for Type I error control. Empirical evaluations on both synthetic and real-world datasets demonstrate that the proposed approach achieves high attribution accuracy and strong robustness in identifying the components responsible for the change.
To address the degradation of prediction probability calibration in deployed image classification models due to concept drift, this paper proposes an online calibration monitoring method that requires no access to model internals—only predicted probabilities and ground-truth labels. Our approach introduces, for the first time, a Cumulative Sum (CUSUM) control chart with dynamic control limits into calibration monitoring. It computes cumulative deviations of calibration error over time and adaptively adjusts detection thresholds to enable early warning of calibration loss. Compared to static-threshold methods, our framework significantly enhances sensitivity to temporal distribution shifts and accelerates response to emerging miscalibration. We validate its effectiveness and robustness across multiple image classification benchmarks under diverse concept drift scenarios. The proposed method establishes a scalable, black-box-compatible paradigm for trustworthy model deployment, enabling continuous, lightweight calibration assessment without architectural or training modifications.
This work addresses the limitation of existing methods that can either detect out-of-distribution samples or quantify uncertainty but struggle to pinpoint the specific causes of model failure. The authors propose a self-diagnosing model that jointly learns structured failure attribution signals alongside its primary predictions, extending scalar uncertainty into an attribution vector capable of distinguishing among four distinct failure modes: covariate shift, semantic shift, noisy corruption, and adversarial perturbations. The attribution vector is generated by a neural network and constrained via a consistency regularization term that aligns uncertainty estimates with attribution predictions. To evaluate the approach, the authors construct a benchmark dataset incorporating predefined shift mechanisms. Experimental results demonstrate that the method not only effectively detects anomalies but also accurately attributes their underlying failure types, significantly enhancing model interpretability and robustness.
Existing ML monitoring systems detect data distribution shifts but struggle to automatically identify their root causes—such as data defects, software failures, or genuine concept drift—relying instead on manual root-cause analysis. To address this, we propose a causal graph framework for ML systems, the first to integrate causal modeling with hierarchical system abstraction. It constructs a multi-granularity dependency graph spanning environmental variables, data pipelines, and model components. By analyzing propagation paths, the framework enables traceable attribution of distribution shifts back to their origins. This systematically establishes a mapping between observed distributional changes and their underlying causes, advancing ML monitoring from merely detecting *whether* a shift occurs to explaining *why* it occurs. The framework provides both theoretical foundations and a practical technical pathway toward interpretable, traceable, and automated ML monitoring. (149 words)