non-iid data handling

Designs, builds, and analyzes methods, evaluation protocols, and data-processing or training procedures that detect, measure, and address violations of the independent-and-identically-distributed (i.i.d.) assumption in datasets. This includes techniques for identifying distribution shifts and dependencies (e.g., covariate/label shift, temporal or spatial correlation, sample heterogeneity, concept drift), and for mitigating their effects via robust training, reweighting/adaptation, specialized evaluation, or other corrective strategies.

non-iiddatahandling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.74
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

When the Past Misleads: Rethinking Training Data Expansion Under Temporal Distribution Shifts

Aug 31, 2025
CY
Chengyuan Yao
🏛️ Columbia University | University of Michigan | Cornell University

This study investigates the impact of expanding the historical training window on predictive performance and algorithmic fairness under temporal distribution shift—including covariate and concept drift. Using simulation experiments and empirical student retention prediction across multi-institution, multi-year educational datasets, we find that the common assumption “more data is better” fails under concept-drift-dominant regimes: extending the training window degrades overall accuracy and exacerbates fairness disparities across sociodemographic groups—particularly when marginalized subpopulations experience heterogeneous concept drift, leading to nonlinear bias amplification. Our key contributions are: (i) identifying concept drift—not covariate drift—as the primary driver of performance degradation; and (ii) the first systematic characterization of its nonlinear, fairness-amplifying mechanism. These findings provide both theoretical grounding and practical guidance for model update strategies in dynamic, real-world deployment environments.

Challenging the assumption that more historical data always improves model outcomesExamining how expanding historical training data affects model performance under temporal shiftsInvestigating the impact of covariate and concept shifts on predictive model fairness

Rethinking Distribution Shifts: Empirical Analysis and Inductive Modeling for Tabular Data

Jul 11, 2023
JL
Jiashuo Liu
🏛️ Tsinghua University | Columbia University

Distribution shift in tabular data lacks empirical grounding, and existing robust learning methods—such as distributionally robust optimization (DRO)—frequently fail in real-world settings. Prior work predominantly assumes covariate shift (X-shift), whereas empirical analysis reveals that label-conditional shift (Y|X-shift) is both more prevalent and dominant in practice. Method: The authors construct a comprehensive benchmark platform spanning five real-world tabular datasets and 60,000 experimental configurations, enabling large-scale ablation studies on robustness. Contribution/Results: (1) Theory-driven methods like DRO offer no consistent advantage over empirical risk minimization (ERM); (2) implementation details—including model selection, hyperparameter tuning, and imbalance handling—exert far greater influence on robustness than the design of ambiguity sets; (3) a data-driven, inductive paradigm should supplant prior structural assumptions. These findings challenge the conventional paradigm in distributionally robust learning and provide reproducible, empirically grounded guidance for co-optimizing algorithms and data.

Analyzing real-world distribution shifts in tabular datasetsEvaluating robust algorithms' performance against empirical shiftsIdentifying implementation factors affecting distributionally robust optimization

This study addresses the bottleneck of concept drift detection in large-scale e-commerce data characterized by hundreds of millions of rows and high-dimensional features. Leveraging Apache Spark, it evaluates multi-column two-sample drift detection methods, with particular emphasis on the scalability of a distributed Maximum Mean Discrepancy (MMD) algorithm integrated with Random Fourier Features. A synthetic injection benchmark comprising 137.5 million rows is constructed, revealing a critical limitation wherein the Kolmogorov-Smirnov test fails due to statistical saturation in identifier-like columns. Experimental results demonstrate that under strong drift conditions, the proposed method achieves a Pearson correlation coefficient of 0.940, a true positive rate of 80.4%, and a false positive rate of merely 3.2%. However, its sensitivity remains constrained in weak drift scenarios.

concept driftdrift detectionhigh-cardinality features

Automatic dataset shift identification to support root cause analysis of AI performance drift

Nov 12, 2024
MR
Mélanie Roschewitz
🏛️ Imperial College London

In AI deployment for medical imaging, data distribution shifts frequently cause abrupt performance degradation and increased misdiagnosis risk. Existing methods can only detect the presence of shift but fail to identify its specific type—e.g., covariate shift, prior (concept) shift, or compound shift—hindering root-cause analysis and targeted mitigation. This paper proposes the first unsupervised framework for data shift type identification. It introduces a novel joint shift detection mechanism that synergistically leverages self-supervised encoder representations and task-model outputs. By integrating feature distribution comparison, unsupervised clustering, and multimodal image modeling, the method achieves high-accuracy shift-type discrimination across three major imaging modalities—chest X-ray, mammography, and fundus photography—and five realistic shift scenarios. Evaluated on four large public medical imaging datasets, it significantly enhances the robustness and interpretability of clinical AI systems.

Distinguish prevalence, covariate, and mixed shifts unsupervisedIdentify diverse dataset shifts in medical imaging AIImprove shift detection using self-supervised encoders

Latest Papers

What's happening recently
View more

When predictive models are deployed in new environments, their performance often degrades due to covariate shift and selective labeling, which jointly obscure accurate assessment of the true target risk. This work proposes an unbiased risk estimation method that integrates double machine learning with influence functions to simultaneously address both sources of bias for the first time. The approach is model-agnostic and compatible with general loss functions, constructing a corrected target risk estimator via nonparametric and plug-in techniques. Experiments on eICU electronic health record data demonstrate that the proposed method significantly outperforms baselines that handle only one type of bias or naively combine existing approaches, yielding more accurate tracking of the true target risk.

covariate shiftdistribution shiftmodel evaluation

This study addresses a critical limitation in existing design-based simulations used to evaluate inference methods, which often overstate bias induced by spatial correlation due to unrealistic data-generating mechanisms. In particular, share-shift designs that fix outcomes and resample shocks conflate true treatment effects with error dependence structures, leading to misleading assessments. To remedy this, the paper proposes an improved simulation framework that more accurately models error dependence and avoids spurious entanglement between treatment effects and error terms, thereby better approximating real-world data-generating processes. Integrating resampling techniques with share-shift analysis, the proposed approach substantially enhances the reliability of inference evaluation across multiple empirical applications, underscoring the essential role of aligning simulation designs with genuine underlying mechanisms for valid inference assessment.

data-generating processdesign-based simulationsinference validity

This study addresses the performance degradation of machine learning classifiers in dynamic environments caused by concept drift, a phenomenon inadequately captured by conventional evaluation methods that overlook causal dependencies in data, leading to distorted assessments. To overcome this limitation, the authors propose a digital twin framework grounded in Structural Causal Models (SCMs), which, for the first time, leverages SCMs to simulate realistic causal drift. By applying parametric causal interventions, the framework stress-tests classifiers while preserving the underlying structure of the data-generating mechanism. This approach transcends the constraints of traditional statistical or correlation-based evaluations. Experimental results on the OSMH dataset demonstrate that the method effectively uncovers classifier vulnerabilities that remain undetected by standard monitoring techniques.

causal dependenciesclassifier robustnessconcept drift

This study addresses the scarcity of tabular data under covariate shift and the instability of conventional augmentation methods caused by fitting outdated source distributions. To this end, we propose the IGDPR framework, which pioneers the integration of invariant potential functions into the diffusion sampling process to align stable decision boundaries and generate task-relevant samples. Furthermore, prototype clustering and reweighting strategies are incorporated to filter noise and assess sample reliability, thereby mitigating overfitting. Experimental results demonstrate that the proposed framework significantly improves synthetic data quality in real-world scenarios, effectively enhancing model robustness and generalization to unseen environments.

Covariate ShiftData AugmentationDistribution Drift

本文解决了机器学习中分布偏移下的泛化问题,通过引入γ*-概念偏移和熵最优传输方法,提出了统一的误差界,并开发了能够估计这些偏移的算法。

concept shiftcovariate shiftdistribution shift

Hot Scholars

AM

Adnan Mahmood

School of Computing, Faculty of Science and Engineering, Macquarie University
Internet of ThingsInternet of VehiclesTrust ManagementSoftware Defined Networking
RB

Rajkumar Buyya

School of Computing and Information Systems, The Uni of Melbourne; Fellow of IEEE & Academia Europea
Cloud ComputingData CentersEdge ComputingInternet of Things
TN

Takayuki Nishio

Associate Professor, School of Engineering, Tokyo Tech
Wireless communicationsMachine LearningmmWaveMobile Computing
WN

Wei Ni

FIEEE, AAIA Fellow, Senior Principal Scientist & Conjoint Professor, CSIRO/UNSW
6G security and privacyconnected and trusted intelligenceapplied AI/ML
ZC

Zhipeng Cai

Professor, IEEE Fellow, DMACM, Georgia State University
Internet of ThingsPrivacyAlgorithmBig Data