asymmetric class-aware augmentation

Designs and implements data-augmentation pipelines that apply transformations asymmetrically according to class identity or prevalence—e.g., generate class-conditional synthetic or simulated examples and targeted lexical variants for minority or fault classes while minimizing or withholding augmentation for abundant/normal classes. This includes devising class-aware augmentation schedules and selection policies (round-robin, asymmetry-aware) to improve classifier training under data scarcity.

asymmetricclass-awareaugmentation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.16
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

A Survey on Small Sample Imbalance Problem: Metrics, Feature Analysis, and Solutions

Apr 21, 2025
SZ
Shuxian Zhao
🏛️ Southeast University | City University of Hong Kong | Nanyang Technological University | Wuhan University | University of Macau | The Hong Kong University of Science and Technology | Purple Mountain Laboratories

The few-shot and imbalanced (S&I) learning problem suffers from severe generalization degradation and low interpretability due to scarce samples, extreme class imbalance, and ambiguous inter-class feature distributions. This paper proposes the first systematic analytical framework tailored to S&I learning, advocating that quantitative characterization of data properties—such as imbalance ratio and geometric complexity—must precede algorithmic design. The framework unifies multi-dimensional imbalance metrics, data complexity analysis, resampling strategies, classifier adaptation mechanisms, and an interpretable evaluation benchmark. Empirical evaluation on binary and multi-class extreme imbalance benchmarks reveals that classifier selection exerts significantly greater impact on performance than resampling improvements—exposing a fundamental flaw in prevailing heuristic-driven approaches. Our work establishes a theory-guided analytical paradigm and practical design principles for S&I learning, advancing both methodological rigor and empirical reproducibility.

Addressing poor model performance in small imbalanced datasetsAnalyzing indistinct inter-class feature distributions for classificationEvaluating effectiveness of resampling versus classifier improvements

Must-Read Papers

Most classic and influential ideas
View more

This study investigates the efficacy limits and optimal scale of synthetic data augmentation in class-imbalanced learning. By establishing a unified statistical learning framework grounded in balanced population risk analysis, the work reveals that augmentation benefits model performance only under “local asymmetry” conditions, and that the optimal number of synthetic minority samples depends critically on the generator’s accuracy and bias direction—challenging the conventional assumption that perfect class balance is inherently optimal. To address this, the authors propose a validation loss–based strategy for tuning the synthetic sample size (VTSS). Both theoretical analysis and empirical experiments demonstrate that ill-conceived augmentation can degrade performance, whereas VTSS reliably identifies the optimal augmentation scale, with consistent validation on both simulated data and real-world sepsis prediction tasks.

class imbalanceimbalanced learningminority class

This study addresses the high computational costs of data augmentation ensembles and their inefficiency in leveraging task symmetries by proposing Stochastic Weight Averaging (SWA) as a replacement for repetitive ensembling. Through approximation analysis via the Ornstein-Uhlenbeck process, we reveal that SWA enhances model equivariance beyond conventional performance gains in the infinite-width limit. Experiments on visual and graph classification tasks demonstrate the method’s superiority across both discrete and continuous symmetries. These findings validate SWA as an effective alternative to traditional ensembling, providing new theoretical foundations and a practical paradigm for efficiently exploiting data augmentation. This work thus bridges the gap between computational efficiency and symmetry-aware learning, offering significant implications for scalable representation learning in structured domains.

Data AugmentationDeep EnsemblesStochastic Weight Averaging

This study investigates the mechanism by which synthetic data augmentation improves score-based classification performance—measured by metrics such as AUROC and AUPRC—in class-imbalanced settings. By developing a theoretical framework that disentangles the effects of augmentation on effective class weighting and distributional bias, and integrating tools from statistical learning theory, minimax analysis, and finite-sample error decomposition, the work establishes that under correctly specified models, augmentation solely reduces variance without improving overall performance. However, under model misspecification, it can mitigate ranking errors by correcting class imbalance. The analysis yields novel minimax lower bounds, which are corroborated through simulation experiments.

class imbalancedistributional discrepancyimbalanced classification

Beyond Rebalancing: Benchmarking Binary Classifiers Under Class Imbalance Without Rebalancing Techniques

Sep 09, 2025
AN
Ali Nawaz
🏛️ United Arab Emirates University | American University of the Middle East

In critical domains such as medical diagnosis, standard binary classifier evaluation under severe class imbalance often fails to reflect real-world robustness, especially when rebalancing techniques are inadmissible. Method: We propose a rebalancing-free robustness evaluation framework that synthesizes complex decision boundaries and adopts few-shot minority-class settings to emulate realistic extreme imbalance. We systematically benchmark TabPFN, ensemble boosting, one-class classification (OCC), and classical sampling methods across multiple real-world and synthetic datasets. Results: Traditional models exhibit significant performance degradation as minority-class prevalence decreases and data complexity increases; in contrast, TabPFN and ensemble methods demonstrate superior generalization and stability. This work is the first to reveal intrinsic robustness disparities among diverse models under unrebalanced conditions within a unified evaluation framework, establishing a new benchmark for imbalanced learning and offering actionable insights for practical deployment.

Assessing classifier robustness with reduced minority class sizesEvaluating binary classifiers without rebalancing under class imbalanceExploring performance across varying data complexities and imbalance scenarios

To address the weak modeling capability of generative models for minority classes in imbalanced tabular classification, this paper proposes a ternary label reconstruction paradigm: extending the original binary labels into “majority class,” “minority class,” and “overlap class” to explicitly model the distributional overlap region between classes. This approach requires no architectural modification to the generative model—only label preprocessing—yet consistently improves minority-class synthesis quality across multiple state-of-the-art generative models, including diffusion models and GAN-based hybrid architectures. Furthermore, an overlap-class removal strategy is introduced to refine downstream classification performance. Extensive experiments across four real-world tabular datasets, five classifiers, and five generative models demonstrate significant and consistent gains in both minority-class sample fidelity and classification accuracy. The method is notably simple, broadly applicable across diverse generative frameworks, and empirically effective.

Addressing class imbalance in tabular data classificationEnhancing classifier accuracy using synthetic dataImproving synthetic data quality for minority classes

Latest Papers

What's happening recently
View more

This work investigates whether partial data augmentation can statistically match the generalization performance and sample complexity of full-group augmentation under computational constraints. By leveraging Fourier analysis and finite group representation theory, the authors establish a unified theoretical framework that, for the first time, characterizes the conditions under which partial and full augmentations are statistically equivalent from a frequency-domain perspective. The main contributions include proving that when the augmented subset is sufficiently large, partial augmentation achieves the same minimax optimal rate as full augmentation, while also demonstrating that exact symmetry—leading to perfect invariance—can only be realized through averaging over the entire group. Consequently, the study delineates the theoretical limits of approximate symmetry and establishes an impossibility result showing that no proper subgroup can yield exact invariance.

computational feasibilitydata augmentationgroup invariance

Bias-Corrected Data Synthesis for Imbalanced Learning

Oct 29, 2025
PL
Pengfei Lyu
🏛️ Duke University | Rutgers University

In class-imbalanced classification, synthesizing minority-class samples often introduces distributional bias, leading to model overfitting and degraded generalization. To address this, we propose a novel framework that estimates and corrects synthesis-induced bias using distributional information from the majority class. Unlike conventional approaches assuming synthesized samples follow the true minority-class distribution, our method leverages structural consistency in majority-class features to construct a provably consistent bias estimator, coupled with dynamic error calibration during training. Theoretically, we derive bounds on the bias estimation error and provide guarantees on improved prediction accuracy. Empirically, extensive experiments on benchmark datasets—including MNIST—demonstrate significant gains in F1-score, AUC, and robustness against label noise. Moreover, the framework naturally extends to multi-task learning and causal inference settings, offering broad applicability without architectural modification.

Correcting bias in synthetic data for imbalanced classification problemsExtending bias correction to multi-task learning and causal inferenceImproving prediction accuracy by mitigating synthetic data adverse effects

This work addresses the severe overfitting that autoregressive language models exhibit under data-constrained yet compute-rich pretraining regimes, where repeated training epochs on a fixed corpus degrade generalization. To mitigate this, the authors propose data augmentation as a regularization mechanism enabling efficient pretraining for hundreds of epochs on static datasets. Three orthogonal augmentation strategies are introduced: token-level noise (e.g., random token replacement), sequence reordering (e.g., right-to-left prediction and infilling), and target-shifted prediction (e.g., forecasting future tokens). Empirical results demonstrate that each strategy effectively reduces validation loss, with random token replacement yielding the strongest individual gains. Combining these augmentations further lowers validation loss, substantially delaying overfitting and enhancing training efficiency.

autoregressive pretrainingdata augmentationdata-constrained pretraining

Current evaluations of bias in code generation are largely confined to simple conditional statements, failing to capture the subtle biases present in real-world programming contexts. This work proposes a systematic evaluation framework grounded in machine learning pipelines, with a specific focus on the introduction of sensitive attributes during the feature selection stage. By testing both code-specific and general-purpose large language models across diverse prompts and complexity levels, the study reveals— for the first time—that existing assessment methods substantially underestimate real-world bias risks: 87.7% of generated ML pipelines incorporate sensitive attributes, markedly higher than the 59.2% detected using conditional-statement-based tests. This discrepancy persists robustly across multiple bias mitigation strategies, thereby challenging the validity of current bias evaluation paradigms.

bias evaluationcode generation biaslarge language models

Hot Scholars

TL

Tianchen Liu

University of Maryland, College Park
OptimizationState EstimationControlCo-design
SB

Sophia Bano

Assistant Professor in Robotics and AI, University College London
Computer VisionSurgical Data ScienceSurgical RoboticsComputer-assisted Intervention
YH

Yuenan Hou

Shanghai AI Laboratory
Autonomous DrivingEmbodied AIEfficient Learning
QZ

Qingfu Zhang

Chair Professor, FIEEE, City University of Hong Kong
evolutionary computationmultiobjective optimizationcomputational intelligence