🤖 AI Summary
To address the weak modeling capability of generative models for minority classes in imbalanced tabular classification, this paper proposes a ternary label reconstruction paradigm: extending the original binary labels into “majority class,” “minority class,” and “overlap class” to explicitly model the distributional overlap region between classes. This approach requires no architectural modification to the generative model—only label preprocessing—yet consistently improves minority-class synthesis quality across multiple state-of-the-art generative models, including diffusion models and GAN-based hybrid architectures. Furthermore, an overlap-class removal strategy is introduced to refine downstream classification performance. Extensive experiments across four real-world tabular datasets, five classifiers, and five generative models demonstrate significant and consistent gains in both minority-class sample fidelity and classification accuracy. The method is notably simple, broadly applicable across diverse generative frameworks, and empirically effective.
📝 Abstract
Handling imbalance in class distribution when building a classifier over tabular data has been a problem of long-standing interest. One popular approach is augmenting the training dataset with synthetically generated data. While classical augmentation techniques were limited to linear interpolation of existing minority class examples, recently higher capacity deep generative models are providing greater promise. However, handling of imbalance in class distribution when building a deep generative model is also a challenging problem, that has not been studied as extensively as imbalanced classifier model training. We show that state-of-the-art deep generative models yield significantly lower-quality minority examples than majority examples. %In this paper, we start with the observation that imbalanced data training of generative models trained imbalanced dataset which under-represent the minority class. We propose a novel technique of converting the binary class labels to ternary class labels by introducing a class for the region where minority and majority distributions overlap. We show that just this pre-processing of the training set, significantly improves the quality of data generated spanning several state-of-the-art diffusion and GAN-based models. While training the classifier using synthetic data, we remove the overlap class from the training data and justify the reasons behind the enhanced accuracy. We perform extensive experiments on four real-life datasets, five different classifiers, and five generative models demonstrating that our method enhances not only the synthesizer performance of state-of-the-art models but also the classifier performance.