distribution shift detection

Designs, builds, and evaluates algorithms, tests, and monitoring systems that detect when the input or feature distribution seen by a model or system deviates from an expected or training distribution (including out‑of‑distribution examples and change‑point detection). Develops and applies statistical tests, scoring functions, evaluation metrics, severity/uncertainty estimation, root‑cause analysis, and operational handling policies (alerts, thresholds, retraining or adaptation triggers) to analyze, validate, and respond to observed distribution shifts.

distributionshiftdetection

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.4
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$216K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

On Continuous Monitoring of Risk Violations under Unknown Shift

Jun 19, 2025
AT
Alexander Timans
🏛️ UvA-Bosch Delta Lab | University of Amsterdam | Johns Hopkins University

Real-world machine learning systems operate under dynamic data distribution shifts, rendering conventional risk control methods—predicated on static distributional assumptions—ineffective and incapable of online monitoring for decision risk violations. To address this, we propose the first sequential testing framework grounded in the “betting” paradigm, which strictly controls the false alarm rate (≤ α) under arbitrary, unknown distributional shifts—without assuming prior knowledge of drift type or underlying distributions. Our approach unifies betting-based hypothesis testing, risk-bound modeling, and online streaming statistical inference to enable real-time, robust monitoring of model risk. Extensive experiments demonstrate that the method achieves high sensitivity in detecting risk violations across diverse drift scenarios, while simultaneously delivering rigorous statistical guarantees in both anomaly detection and conformal prediction tasks.

Detect bounded risk violations with false alarm controlEnsure reliability under unknown distribution shiftsMonitor risk violations in dynamic data streams

To address model performance degradation caused by data distribution drift and the reliance of existing MLOps retraining pipelines on manual intervention, this paper proposes an automated, adaptive neural network retraining framework. Methodologically, it introduces a novel multi-criteria joint drift detection mechanism—integrating statistical metrics including the Kolmogorov–Smirnov test, Population Stability Index (PSI), and Classifier-Driven (CD) drift detection—combined with online monitoring and lightweight scheduling to dynamically trigger end-to-end retraining upon significant drift. The framework is implemented using a cloud-native architecture for scalable and efficient deployment. Evaluated on multiple benchmark datasets, the proposed approach improves classification accuracy by 12.3%–18.7%, reduces inference latency by 41%, and cuts computational resource consumption by 53%, compared to conventional periodic or single-threshold retraining strategies. These gains significantly enhance model freshness and operational cost-efficiency.

Automates MLOps for retraining classifiers due to data driftImproves accuracy and robustness in dynamic real-world settingsUses multi-criteria detection to trigger updates only when needed

Multiple Distribution Shift -- Aerial (MDS-A): A Dataset for Test-Time Error Detection and Model Adaptation

Feb 18, 2025
NN
Noel Ngu
🏛️ Arizona State University | Universidad Nacional del Sur | U.S. Department of Defense | United States Military Academy

This work addresses the severe performance degradation and erroneous detection of aerial vision models under weather-induced distribution shifts. To this end, we introduce MDS-A—the first multi-distribution-shift benchmark for aerial imagery—featuring high-fidelity synthetic training data generated in Unreal Engine under six controlled meteorological conditions, a mixed-weather test set, and comprehensive annotations. We propose EDR (Error Detection and Recovery), a knowledge-driven framework enabling test-time uncertainty modeling and lightweight self-adaptation. MDS-A is the first benchmark to support fine-grained out-of-distribution (OOD) attribution analysis and standardized evaluation across multidimensional weather shifts. Experiments show that mainstream YOLOv5/v8 models suffer 32–68% mAP drops across weather domains; with EDR, erroneous detection accuracy reaches 89.7%, and online model adaptation is effectively triggered.

Addresses performance degradation due to distribution shiftsEvaluates models under varied simulated weather conditionsIntroduces MDS-A dataset for error detection and adaptation

This work addresses the limitation of existing safety-critical systems, which typically evaluate only predictive accuracy while lacking rigorous validation of the overall calibration of predicted probability distributions. To bridge this gap, the authors propose a modular calibration testing framework that decouples the calibration process into four interchangeable components: data model, scoring rule, hypothesis formulation, and statistical test procedure. Built upon formal statistical hypothesis testing, the framework provides a single accept/reject decision for the entire predictive distribution. Crucially, it rejects only overly confident predictions while tolerating reasonable deviations, thereby balancing practicality with flexibility. Empirical evaluations on weather forecasting and robotic pose estimation tasks demonstrate that the framework effectively supports reliable deployment in safety-critical applications.

calibrationdistributional validationprobabilistic forecasting

In industrial quality inspection, anomaly detection suffers from poor robustness due to high noise levels and sparse defective samples. To address this, we propose Iterative Refinement of Pseudo-labels (IRP), a self-supervised method that alternately evaluates sample credibility and removes misleading instances under feature-space consistency constraints—effectively purifying the training set dynamically without human annotations and generating high-fidelity self-supervised signals. IRP introduces the novel paradigm of “iterative data refinement,” significantly enhancing model robustness against label noise and cross-domain generalization capability. Evaluated on KSDD2 and MVTec AD benchmarks, IRP consistently outperforms existing unsupervised and self-supervised methods. Notably, under high-noise conditions, it achieves substantial improvements in detection accuracy and reduces false positive rates by over 25%.

Enhances defect detection accuracy in industrial quality control.Improves model performance by removing misleading data points.Outperforms traditional models in noisy industrial environments.

Latest Papers

What's happening recently
View more

This work addresses the challenge of accurately attributing detected change points in multivariate time series to specific subsets of variables. The authors propose a post-hoc, nonparametric testing framework that, after an offline change point has been identified, determines whether the change occurs in one of two pre-specified coordinate blocks or in both. Built upon two-sample nonparametric hypothesis testing, the method offers rigorous theoretical guarantees for Type I error control. Empirical evaluations on both synthetic and real-world datasets demonstrate that the proposed approach achieves high attribution accuracy and strong robustness in identifying the components responsible for the change.

change-point detectioncomponent attributionmultivariate time series

To address the degradation of prediction probability calibration in deployed image classification models due to concept drift, this paper proposes an online calibration monitoring method that requires no access to model internals—only predicted probabilities and ground-truth labels. Our approach introduces, for the first time, a Cumulative Sum (CUSUM) control chart with dynamic control limits into calibration monitoring. It computes cumulative deviations of calibration error over time and adaptively adjusts detection thresholds to enable early warning of calibration loss. Compared to static-threshold methods, our framework significantly enhances sensitivity to temporal distribution shifts and accelerates response to emerging miscalibration. We validate its effectiveness and robustness across multiple image classification benchmarks under diverse concept drift scenarios. The proposed method establishes a scalable, black-box-compatible paradigm for trustworthy model deployment, enabling continuous, lightweight calibration assessment without architectural or training modifications.

Assessing prediction calibration without accessing the underlying ML modelDetecting concept drift affecting image classification model performanceMonitoring calibration loss in probability forecasts over time

This work addresses the limitation of existing methods that can either detect out-of-distribution samples or quantify uncertainty but struggle to pinpoint the specific causes of model failure. The authors propose a self-diagnosing model that jointly learns structured failure attribution signals alongside its primary predictions, extending scalar uncertainty into an attribution vector capable of distinguishing among four distinct failure modes: covariate shift, semantic shift, noisy corruption, and adversarial perturbations. The attribution vector is generated by a neural network and constrained via a consistency regularization term that aligns uncertainty estimates with attribution predictions. To evaluate the approach, the authors construct a benchmark dataset incorporating predefined shift mechanisms. Experimental results demonstrate that the method not only effectively detects anomalies but also accurately attributes their underlying failure types, significantly enhancing model interpretability and robustness.

distribution shiftfailure attributionmodel robustness

Tracing Distribution Shifts with Causal System Maps

Oct 27, 2025
JL
Joran Leest
🏛️ Vrije Universiteit | Universita’ degli Studi di Milano-Bicocca

Existing ML monitoring systems detect data distribution shifts but struggle to automatically identify their root causes—such as data defects, software failures, or genuine concept drift—relying instead on manual root-cause analysis. To address this, we propose a causal graph framework for ML systems, the first to integrate causal modeling with hierarchical system abstraction. It constructs a multi-granularity dependency graph spanning environmental variables, data pipelines, and model components. By analyzing propagation paths, the framework enables traceable attribution of distribution shifts back to their origins. This systematically establishes a mapping between observed distributional changes and their underlying causes, advancing ML monitoring from merely detecting *whether* a shift occurs to explaining *why* it occurs. The framework provides both theoretical foundations and a practical technical pathway toward interpretable, traceable, and automated ML monitoring. (149 words)

Automating root-cause analysis for data quality issuesDetecting causes of ML system distribution shiftsMapping causal propagation paths in ML systems

Hot Scholars

AN

Amin Nikanjam

Staff Researcher, Huawei Canada
LLMsMachine Learning Systems EngineeringSE4ML
FK

Foutse Khomh

NSERC Arthur B. McDonald Fellow, CRC Tier 1, Canada CIFAR AI Chair, FRQ-IVADO Chair, Full Professor
Software engineeringMachine learning systems engineeringMining software repositoriesReverse
YL

Yuxuan Liang

Assistant Professor, Hong Kong University of Science and Technology (Guangzhou)
Spatio-Temporal Data MiningUrban ComputingUrban AIFoundation Models
SP

Shirui Pan

Professor, ARC Future Fellow, FQA, Director of TrustAGI Lab, Griffith University
Data MiningMachine LearningGraph Neural NetworksTrustworthy AI
ZZ

Zhuosheng Zhang

Assistant Professor at Shanghai Jiao Tong University
Natural Language ProcessingLarge Language ModelsReasoningAI Safety