🤖 AI Summary
To address model bias toward majority classes in imbalanced classification, this paper proposes an end-to-end trainable deep oversampling framework. The method employs a parameterized transformation to map majority-class samples into the minority-class distribution space. It innovatively integrates Maximum Mean Discrepancy (MMD) for global distribution alignment and incorporates triplet loss to guide synthetic sample generation toward challenging regions near the decision boundary, thereby significantly enhancing boundary-awareness. Extensive experiments across 29 standard benchmark datasets demonstrate that the proposed approach consistently outperforms conventional resampling techniques and generative baselines across key metrics—including AUROC, G-mean, F1-score, and Matthews Correlation Coefficient (MCC)—validating its robustness and effectiveness in mitigating class imbalance.
📝 Abstract
Class imbalance in supervised classification often degrades model performance by biasing predictions toward the majority class, particularly in critical applications such as medical diagnosis and fraud detection. Traditional oversampling techniques, including SMOTE and its variants, generate synthetic minority samples via local interpolation but fail to capture global data distributions in high-dimensional spaces. Deep generative models based on GANs offer richer distribution modeling yet suffer from training instability and mode collapse under severe imbalance. To overcome these limitations, we introduce an oversampling framework that learns a parametric transformation to map majority samples into the minority distribution. Our approach minimizes the maximum mean discrepancy (MMD) between transformed and true minority samples for global alignment, and incorporates a triplet loss regularizer to enforce boundary awareness by guiding synthesized samples toward challenging borderline regions. We evaluate our method on 29 synthetic and real-world datasets, demonstrating consistent improvements over classical and generative baselines in AUROC, G-mean, F1-score, and MCC. These results confirm the robustness, computational efficiency, and practical utility of the proposed framework for imbalanced classification tasks.