quality-controlled oversampling

Designs and implements oversampling algorithms that produce synthetic minority-class examples while explicitly assessing and enforcing sample quality; this includes mechanisms to generate candidate synthetics, score or estimate their reliability, apply geometry-adaptive interpolation (SMOTE variants), and perform best-of-k candidate selection or replacement of low-quality synthetics with duplicates. The skill covers building pipelines to create, select, and integrate quality-controlled synthetic samples into training sets and analyzing their effect on class balance and model performance.

quality-controlledoversampling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.07
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the limitation of conventional oversampling methods like SMOTE, which often generate low-quality synthetic samples in noisy or class-overlapping regions under class imbalance. To overcome this, the authors propose a quality-controllable oversampling framework that evaluates the reliability of minority-class instances using a composite neighborhood credibility score—integrating local density, safety level, and isolation from majority classes. High-quality synthetic samples are then generated via an IPQ-guided Best-of-K selection strategy. Furthermore, the method adaptively adjusts interpolation ranges and selection criteria based on local data geometry, reverting to simple replication in low-purity regions to enhance robustness. Experimental results across 30 imbalanced datasets demonstrate that the proposed approach consistently outperforms existing oversampling techniques in terms of AUC-ROC and Macro F1, with particularly notable gains under moderate to severe class imbalance.

class imbalanceclass overlapnoise

Concentration and excess risk bounds for imbalanced classification with synthetic oversampling

Oct 23, 2025
TA
Touqeer Ahmad
🏛️ Univ Rennes | Ensai | CNRS | CREST—UMR 9194 | Univ Angers | LAREMA | SFR MATHSTIC

This paper addresses the lack of theoretical foundations for synthetic oversampling methods such as SMOTE in imbalanced classification. We establish, for the first time, a statistical learning theory framework for such methods. By integrating uniform concentration inequalities with nonparametric estimation, we rigorously characterize the convergence between the empirical risk—computed over synthetically augmented data—and the population risk under the true data distribution, and derive a nonparametric excess risk bound for kernel classifiers. Our key contributions are: (1) the first unified concentration bound applicable to SMOTE-like oversampling schemes; (2) a theoretically grounded criterion for joint tuning of oversampling parameters (e.g., neighborhood size, synthesis ratio) and classifier hyperparameters; and (3) empirical validation demonstrating that the derived bound effectively predicts and guides generalization performance in practice.

Analyzing theoretical foundations of SMOTE for imbalanced classificationDeriving concentration bounds between synthetic and true minority distributionsProviding excess risk guarantees for kernel classifiers with synthetic data

Do we need rebalancing strategies? A theoretical and empirical study around SMOTE and its variants

Feb 06, 2024
AS
Abdoulaye Sakho
🏛️ Artefact Research Center | Sorbonne Université

This work investigates the necessity and efficacy of rebalancing strategies—particularly SMOTE and its variants—for imbalanced tabular data, through both theoretical analysis and empirical evaluation. We derive, for the first time, a non-asymptotic upper bound on the density induced by SMOTE, rigorously proving that, under default parameters, it degenerates to mere sample duplication and yields vanishing density near class boundaries—revealing intrinsic “density degradation” and “boundary failure.” Guided by this theory, we propose two novel SMOTE variants. Comprehensive evaluation across 13 benchmark datasets, 10 rebalancing methods (including diffusion-based approaches), and strong baselines such as LightGBM demonstrates that competitive performance is often achievable without rebalancing in realistic scenarios; moreover, as class imbalance intensifies, our variants significantly outperform standard SMOTE and state-of-the-art alternatives.

Evaluation of SMOTE variants vs. state-of-the-art rebalancing methodsImpact of imbalance ratio on rebalancing strategy effectivenessTheoretical analysis of SMOTE's density and asymptotic behavior

This study addresses the issue of model bias toward majority classes in imbalanced classification by formally framing it as a label shift domain adaptation problem between the source distribution (observed data) and the target distribution (balanced evaluation distribution). The authors introduce the concept of “transfer cost” and provide theoretical analysis showing that SMOTE incurs higher transfer cost than random oversampling methods such as Bootstrap in medium- to high-dimensional spaces. Building on this framework, they integrate minority class distribution estimation into data augmentation and empirically demonstrate that random oversampling generally outperforms SMOTE in such settings. These findings offer both theoretical justification and practical guidance for selecting oversampling strategies in imbalanced classification tasks.

classification imbalancelabel shiftoversampling

Synthetic Oversampling: Theory and A Practical Approach Using LLMs to Address Data Imbalance

Jun 05, 2024
RN
Ryumei Nakada
🏛️ Rutgers University | University of California, Berkeley

In imbalanced classification, scarcity of minority-class samples induces model bias and spurious correlations. Method: This paper proposes a novel synthetic oversampling paradigm leveraging large language models (LLMs), establishing the first theoretical framework for synthetic data in imbalanced learning. It rigorously quantifies performance gains, derives scaling laws linking synthetic sample size to model accuracy, and characterizes the capability boundary of Transformers for generating high-fidelity synthetic samples. Contribution/Results: Theoretically, the method provably enhances classification accuracy, robustness, and generalization. Empirically, LLM-generated samples effectively mitigate class bias and outperform conventional resampling techniques (e.g., SMOTE) across multiple benchmarks. This work delivers an interpretable, scalable, and LLM-driven solution for trustworthy imbalanced learning.

Imbalanced ClassificationMachine LearningSpurious Correlations

Latest Papers

What's happening recently
View more

This study addresses the adverse impact of resampling methods—such as SMOTE and random undersampling—on the calibration of tree-based ensemble models under class imbalance. While these techniques improve classification performance, they degrade probability calibration, thereby compromising the reliability of decisions that depend on predicted probabilities. The work systematically evaluates this effect and quantifies, for the first time, that SMOTE increases the expected calibration error (ECE) by an average of 0.009, whereas random undersampling under high imbalance elevates ECE to as much as 0.395. It further demonstrates that standard prior-probability correction is ineffective for SMOTE, necessitating data-driven post-hoc calibration. Experiments show that applying Platt or isotonic regression reduces ECE by up to 66% with negligible AUC degradation (only 0.002), underscoring the necessity and efficacy of post-calibration in imbalanced learning scenarios.

class imbalanceprobability calibrationresampling

To address model bias toward majority classes in imbalanced classification, this paper proposes an end-to-end trainable deep oversampling framework. The method employs a parameterized transformation to map majority-class samples into the minority-class distribution space. It innovatively integrates Maximum Mean Discrepancy (MMD) for global distribution alignment and incorporates triplet loss to guide synthetic sample generation toward challenging regions near the decision boundary, thereby significantly enhancing boundary-awareness. Extensive experiments across 29 standard benchmark datasets demonstrate that the proposed approach consistently outperforms conventional resampling techniques and generative baselines across key metrics—including AUROC, G-mean, F1-score, and Matthews Correlation Coefficient (MCC)—validating its robustness and effectiveness in mitigating class imbalance.

Addresses class imbalance degrading model performance in supervised classificationLearns parametric transformation mapping majority to minority distribution with MMD and triplet lossOvercomes limitations of traditional oversampling and generative models for imbalance

This study investigates the efficacy limits and optimal scale of synthetic data augmentation in class-imbalanced learning. By establishing a unified statistical learning framework grounded in balanced population risk analysis, the work reveals that augmentation benefits model performance only under “local asymmetry” conditions, and that the optimal number of synthetic minority samples depends critically on the generator’s accuracy and bias direction—challenging the conventional assumption that perfect class balance is inherently optimal. To address this, the authors propose a validation loss–based strategy for tuning the synthetic sample size (VTSS). Both theoretical analysis and empirical experiments demonstrate that ill-conceived augmentation can degrade performance, whereas VTSS reliably identifies the optimal augmentation scale, with consistent validation on both simulated data and real-world sepsis prediction tasks.

class imbalanceimbalanced learningminority class

GK-SMOTE: A Hyperparameter-free Noise-Resilient Gaussian KDE-Based Oversampling Approach

Sep 14, 2025
MR
Mahabubur Rahman Miraj
🏛️ Chongqing University | Michigan Technological University

To address the degradation of model performance in imbalanced classification caused by label noise and complex class distributions, this paper proposes a hyperparameter-free, noise-robust density-aware oversampling method. The approach employs Gaussian kernel density estimation (KDE) to adaptively identify high-density “safe” regions and low-density “noisy” or ambiguous regions; synthetic samples are generated exclusively within safe regions. It further integrates a boundary-aware identification strategy into an enhanced SMOTE framework. Its core innovation lies in a density-driven regional discrimination mechanism that inherently avoids noise contamination, thereby significantly improving class separability and model robustness. Extensive experiments on multiple binary-class benchmark datasets demonstrate that the proposed method consistently outperforms state-of-the-art oversampling techniques across key metrics—including Matthews Correlation Coefficient (MCC), balanced accuracy, and Area Under the Precision-Recall Curve (AUPRC)—particularly under realistic noisy conditions.

Addresses imbalanced classification challenges in machine learningHandles label noise and complex data distributions effectivelyImproves classification accuracy without requiring hyperparameter tuning

This paper exposes a critical privacy leakage risk of SMOTE in privacy-sensitive settings: its minority-class oversampling process inadvertently reveals original sensitive records, undetectable by conventional evaluation methods. To demonstrate this vulnerability, we propose two novel adversarial attacks—DistinSMOTE, which exploits geometric feature disparities to distinguish real from synthetic samples, and ReconSMOTE, which achieves high-fidelity reconstruction of original minority-class instances. Leveraging membership inference, distance-based analysis, and geometric modeling—supported by theoretical proofs and extensive experiments—we evaluate both attacks across eight diverse imbalanced datasets. Under typical class-imbalance ratios, both methods achieve near-perfect recall and precision (≈100%). This work provides the first systematic evidence that SMOTE offers no inherent privacy protection, delivering a crucial cautionary insight and establishing a new benchmark for co-designing fairness-aware balancing techniques and privacy-preserving machine learning.

Developing attacks to distinguish and reconstruct recordsExposing privacy leakage in SMOTE data synthesisRevealing disproportionate minority record exposure risks

Hot Scholars

MY

Mohammad Yaqub

Researcher in Biomedical Engineering, Associate professor at MBZUAI
Artificial IntelligenceMedical Image AnalysisMachine LearningDeep learning
LG

Luca Guarnera

University of Catania - IPLab (Image Processing Lab)
Multimedia ForensicsComputer VisionMachine learningPattern recognition
SB

Sebastiano Battiato

Full Professor, University of Catania - IPLab (Image Processing lab) - ICTLab
Computer VisionInformation Forensics and SecurityMultimedia ForensicsMedical Imaging
RC

Rohitash Chandra

UNSW
Bayesian deep learningNeuroevolutionClimate ExtremesLanguage Models
DM

Deyu Meng

Professor, Xi'an Jiaotong University
Machine LearningApplied MathematicsComputer VisionArtificial Intelligence