Score
Designs and implements data-augmentation pipelines that apply transformations asymmetrically according to class identity or prevalence—e.g., generate class-conditional synthetic or simulated examples and targeted lexical variants for minority or fault classes while minimizing or withholding augmentation for abundant/normal classes. This includes devising class-aware augmentation schedules and selection policies (round-robin, asymmetry-aware) to improve classifier training under data scarcity.
Class imbalance severely degrades model discrimination for minority classes, critically hindering deployment in high-stakes domains such as healthcare and finance. This paper systematically surveys over one hundred imbalance mitigation strategies, introducing the first unified taxonomy that integrates generative approaches (e.g., GANs, VAEs) with classical resampling techniques—including SMOTE, neighborhood density estimation, and adaptive threshold-based resampling. We further propose a multidimensional evaluation framework and practical deployment guidelines tailored to real-world constraints. Empirical validation across diverse benchmark tasks demonstrates that the surveyed methods improve minority-class F1-score by 12–35%. Crucially, we identify a novel pathway for jointly optimizing interpretability and generalization—bridging theoretical advances with engineering feasibility. This work provides a comprehensive, actionable foundation for both advancing imbalance learning theory and enabling robust, trustworthy deployment in critical applications.
The few-shot and imbalanced (S&I) learning problem suffers from severe generalization degradation and low interpretability due to scarce samples, extreme class imbalance, and ambiguous inter-class feature distributions. This paper proposes the first systematic analytical framework tailored to S&I learning, advocating that quantitative characterization of data properties—such as imbalance ratio and geometric complexity—must precede algorithmic design. The framework unifies multi-dimensional imbalance metrics, data complexity analysis, resampling strategies, classifier adaptation mechanisms, and an interpretable evaluation benchmark. Empirical evaluation on binary and multi-class extreme imbalance benchmarks reveals that classifier selection exerts significantly greater impact on performance than resampling improvements—exposing a fundamental flaw in prevailing heuristic-driven approaches. Our work establishes a theory-guided analytical paradigm and practical design principles for S&I learning, advancing both methodological rigor and empirical reproducibility.
This study investigates the efficacy limits and optimal scale of synthetic data augmentation in class-imbalanced learning. By establishing a unified statistical learning framework grounded in balanced population risk analysis, the work reveals that augmentation benefits model performance only under “local asymmetry” conditions, and that the optimal number of synthetic minority samples depends critically on the generator’s accuracy and bias direction—challenging the conventional assumption that perfect class balance is inherently optimal. To address this, the authors propose a validation loss–based strategy for tuning the synthetic sample size (VTSS). Both theoretical analysis and empirical experiments demonstrate that ill-conceived augmentation can degrade performance, whereas VTSS reliably identifies the optimal augmentation scale, with consistent validation on both simulated data and real-world sepsis prediction tasks.
This study addresses the high computational costs of data augmentation ensembles and their inefficiency in leveraging task symmetries by proposing Stochastic Weight Averaging (SWA) as a replacement for repetitive ensembling. Through approximation analysis via the Ornstein-Uhlenbeck process, we reveal that SWA enhances model equivariance beyond conventional performance gains in the infinite-width limit. Experiments on visual and graph classification tasks demonstrate the method’s superiority across both discrete and continuous symmetries. These findings validate SWA as an effective alternative to traditional ensembling, providing new theoretical foundations and a practical paradigm for efficiently exploiting data augmentation. This work thus bridges the gap between computational efficiency and symmetry-aware learning, offering significant implications for scalable representation learning in structured domains.
This study investigates the mechanism by which synthetic data augmentation improves score-based classification performance—measured by metrics such as AUROC and AUPRC—in class-imbalanced settings. By developing a theoretical framework that disentangles the effects of augmentation on effective class weighting and distributional bias, and integrating tools from statistical learning theory, minimax analysis, and finite-sample error decomposition, the work establishes that under correctly specified models, augmentation solely reduces variance without improving overall performance. However, under model misspecification, it can mitigate ranking errors by correcting class imbalance. The analysis yields novel minimax lower bounds, which are corroborated through simulation experiments.
In critical domains such as medical diagnosis, standard binary classifier evaluation under severe class imbalance often fails to reflect real-world robustness, especially when rebalancing techniques are inadmissible. Method: We propose a rebalancing-free robustness evaluation framework that synthesizes complex decision boundaries and adopts few-shot minority-class settings to emulate realistic extreme imbalance. We systematically benchmark TabPFN, ensemble boosting, one-class classification (OCC), and classical sampling methods across multiple real-world and synthetic datasets. Results: Traditional models exhibit significant performance degradation as minority-class prevalence decreases and data complexity increases; in contrast, TabPFN and ensemble methods demonstrate superior generalization and stability. This work is the first to reveal intrinsic robustness disparities among diverse models under unrebalanced conditions within a unified evaluation framework, establishing a new benchmark for imbalanced learning and offering actionable insights for practical deployment.
To address the weak modeling capability of generative models for minority classes in imbalanced tabular classification, this paper proposes a ternary label reconstruction paradigm: extending the original binary labels into “majority class,” “minority class,” and “overlap class” to explicitly model the distributional overlap region between classes. This approach requires no architectural modification to the generative model—only label preprocessing—yet consistently improves minority-class synthesis quality across multiple state-of-the-art generative models, including diffusion models and GAN-based hybrid architectures. Furthermore, an overlap-class removal strategy is introduced to refine downstream classification performance. Extensive experiments across four real-world tabular datasets, five classifiers, and five generative models demonstrate significant and consistent gains in both minority-class sample fidelity and classification accuracy. The method is notably simple, broadly applicable across diverse generative frameworks, and empirically effective.
This work investigates whether partial data augmentation can statistically match the generalization performance and sample complexity of full-group augmentation under computational constraints. By leveraging Fourier analysis and finite group representation theory, the authors establish a unified theoretical framework that, for the first time, characterizes the conditions under which partial and full augmentations are statistically equivalent from a frequency-domain perspective. The main contributions include proving that when the augmented subset is sufficiently large, partial augmentation achieves the same minimax optimal rate as full augmentation, while also demonstrating that exact symmetry—leading to perfect invariance—can only be realized through averaging over the entire group. Consequently, the study delineates the theoretical limits of approximate symmetry and establishes an impossibility result showing that no proper subgroup can yield exact invariance.
本文通过Wasserstein距离研究生成式数据增强对分类风险的影响,提出基于Rademacher复杂度的一般化边界,评估了不同生成模型在缓解类别不平衡问题上的表现。
In class-imbalanced classification, synthesizing minority-class samples often introduces distributional bias, leading to model overfitting and degraded generalization. To address this, we propose a novel framework that estimates and corrects synthesis-induced bias using distributional information from the majority class. Unlike conventional approaches assuming synthesized samples follow the true minority-class distribution, our method leverages structural consistency in majority-class features to construct a provably consistent bias estimator, coupled with dynamic error calibration during training. Theoretically, we derive bounds on the bias estimation error and provide guarantees on improved prediction accuracy. Empirically, extensive experiments on benchmark datasets—including MNIST—demonstrate significant gains in F1-score, AUC, and robustness against label noise. Moreover, the framework naturally extends to multi-task learning and causal inference settings, offering broad applicability without architectural modification.
This work addresses the severe overfitting that autoregressive language models exhibit under data-constrained yet compute-rich pretraining regimes, where repeated training epochs on a fixed corpus degrade generalization. To mitigate this, the authors propose data augmentation as a regularization mechanism enabling efficient pretraining for hundreds of epochs on static datasets. Three orthogonal augmentation strategies are introduced: token-level noise (e.g., random token replacement), sequence reordering (e.g., right-to-left prediction and infilling), and target-shifted prediction (e.g., forecasting future tokens). Empirical results demonstrate that each strategy effectively reduces validation loss, with random token replacement yielding the strongest individual gains. Combining these augmentations further lowers validation loss, substantially delaying overfitting and enhancing training efficiency.
Current evaluations of bias in code generation are largely confined to simple conditional statements, failing to capture the subtle biases present in real-world programming contexts. This work proposes a systematic evaluation framework grounded in machine learning pipelines, with a specific focus on the introduction of sensitive attributes during the feature selection stage. By testing both code-specific and general-purpose large language models across diverse prompts and complexity levels, the study reveals— for the first time—that existing assessment methods substantially underestimate real-world bias risks: 87.7% of generated ML pipelines incorporate sensitive attributes, markedly higher than the 59.2% detected using conditional-statement-based tests. This discrepancy persists robustly across multiple bias mitigation strategies, thereby challenging the validity of current bias evaluation paradigms.