Rethinking Data Augmentation under Covariate Shift: Invariant-Guided Diffusion and Prototype Reweighting

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the scarcity of tabular data under covariate shift and the instability of conventional augmentation methods caused by fitting outdated source distributions. To this end, we propose the IGDPR framework, which pioneers the integration of invariant potential functions into the diffusion sampling process to align stable decision boundaries and generate task-relevant samples. Furthermore, prototype clustering and reweighting strategies are incorporated to filter noise and assess sample reliability, thereby mitigating overfitting. Experimental results demonstrate that the proposed framework significantly improves synthetic data quality in real-world scenarios, effectively enhancing model robustness and generalization to unseen environments.
📝 Abstract
In many industrial applications, 1) tabular data is scarce and imbalanced and thus requires synthetic expansion; 2) input distributions drift between training and deployment (covariate shift); 3) validation sets often diverge from unseen test environments; or 4) standard generative models simply mimic outdated source distributions. This learning setting limits the stability of standard augmentation and adaptation pipelines. We generalize the task under such setting as the Augmented and Weighted Learning under Covariate Shift problem (AWL-CS). AWL-CS imposes two critical challenges on existing methods: 1) misleading generative guidance where models optimize for source similarity rather than downstream task relevance, and 2) structural instability of distributional density where reweighting mechanisms overfit to noisy validation signals. To tackle these challenges, we propose IGDPR (Invariant-Guided Diffusion with Prototype Reweighting), a unified framework that synergizes stable synthesis and structural adaptation: i) To achieve task-relevant generation, we steer the diffusion sampling process using invariant potentials to ensure synthetic samples align with stable decision boundaries rather than outdated correlations. ii) To ensure stable adaptation, we develop a prototype-based reweighting strategy that assesses sample reliability through structural clusters instead of isolated points, effectively filtering validation noise. Extensive experiments on real data demonstrate our method improves data quality by augmenting the most beneficial data for robust learning.
Problem

Research questions and friction points this paper is trying to address.

Covariate Shift
Data Augmentation
Tabular Data
Distribution Drift
Reweighting
Innovation

Methods, ideas, or system contributions that make the work stand out.

Covariate Shift
Diffusion Model
Invariant-Guided Generation
Prototype Reweighting
Data Augmentation
🔎 Similar Papers
No similar papers found.