imbalanced classification

Designs and implements methods and pipelines to mitigate class-imbalance effects in supervised classification, including sampling strategies (oversampling, SMOTE, stratified splits), loss-based approaches (class weighting, reweighting, focal-loss calibration), and deployment of class-balancing techniques to improve minority-class performance. Analyzes dataset class distributions and training dynamics—such as per-class metrics, gradient contributions, and overfitting to head classes—and selects or tunes rebalancing strategies and hyperparameters to equalize class influence.

imbalancedclassification

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.27
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$204K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

A Review of Machine Learning Techniques in Imbalanced Data and Future Trends

Oct 11, 2023
EJ
Elaheh Jafarigol
🏛️ University of Oklahoma

Class imbalance severely biases model training and undermines evaluation validity across domains. Method: This paper systematically reviews 258 authoritative publications (2003–2023) and proposes the first cross-domain, multi-dimensional taxonomy for imbalance learning—unifying sampling-based methods (e.g., SMOTE, ADASYN), cost-sensitive learning, ensemble techniques (e.g., EasyEnsemble, RUSBoost), deep learning adaptations, and evaluation metrics (F1, G-mean, AUC-PR). Contribution/Results: We construct a full-stack knowledge graph spanning preprocessing, modeling, evaluation, and deployment, and introduce the first evaluation selection guideline tailored to large-scale, real-world imbalanced applications. The framework significantly lowers practical adoption barriers in high-skew domains such as financial risk control and medical diagnosis, while identifying emerging research frontiers—including self-supervised and causal learning integration.

Addressing rare event detection challenges in real-world applicationsProviding guidelines for handling large-scale imbalanced datasetsReviewing machine learning techniques for imbalanced data

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the issue of model bias toward majority classes in imbalanced classification by formally framing it as a label shift domain adaptation problem between the source distribution (observed data) and the target distribution (balanced evaluation distribution). The authors introduce the concept of “transfer cost” and provide theoretical analysis showing that SMOTE incurs higher transfer cost than random oversampling methods such as Bootstrap in medium- to high-dimensional spaces. Building on this framework, they integrate minority class distribution estimation into data augmentation and empirically demonstrate that random oversampling generally outperforms SMOTE in such settings. These findings offer both theoretical justification and practical guidance for selecting oversampling strategies in imbalanced classification tasks.

classification imbalancelabel shiftoversampling

Restoring balance: principled under/oversampling of data for optimal classification

May 15, 2024
EL
Emanuele Loffredo
🏛️ PSL University | Sorbonne University | Université Paris-Cité

Linear classifiers (e.g., SVM) suffer from degraded generalization performance on high-dimensional imbalanced data. Method: We establish a high-dimensional asymptotic theoretical framework and, for the first time, rigorously derive analytical expressions for the generalization error under undersampling and oversampling. Our approach integrates random matrix theory, high-dimensional statistical learning, and unsupervised probabilistic modeling–driven resampling. Contribution/Results: We quantify how resampling efficacy depends on the first- and second-order statistics of the data and the choice of evaluation metric. Crucially, we prove—and empirically verify—that hybrid sampling consistently outperforms either undersampling or oversampling alone. Extensive numerical experiments and evaluations on real-world datasets—including deep neural network features—demonstrate strong agreement between theoretical predictions and empirical results, with substantial improvements in minority-class classification accuracy. This work provides an interpretable, generalizable, and principle-based foundation for data rebalancing in high dimensions.

High-dimensional dataImbalanced datasetsLinear classifiers

Data Balancing Strategies: A Survey of Resampling and Augmentation Methods

May 17, 2025
BY
Behnam Yousefimehr
🏛️ Amirkabir University of Technology

Class imbalance severely degrades model discrimination for minority classes, critically hindering deployment in high-stakes domains such as healthcare and finance. This paper systematically surveys over one hundred imbalance mitigation strategies, introducing the first unified taxonomy that integrates generative approaches (e.g., GANs, VAEs) with classical resampling techniques—including SMOTE, neighborhood density estimation, and adaptive threshold-based resampling. We further propose a multidimensional evaluation framework and practical deployment guidelines tailored to real-world constraints. Empirical validation across diverse benchmark tasks demonstrates that the surveyed methods improve minority-class F1-score by 12–35%. Crucially, we identify a novel pathway for jointly optimizing interpretability and generalization—bridging theoretical advances with engineering feasibility. This work provides a comprehensive, actionable foundation for both advancing imbalance learning theory and enabling robust, trustworthy deployment in critical applications.

Addressing imbalanced data skewing machine learning predictionsExploring resampling methods like SMOTE and undersampling techniquesReviewing advanced generative models for synthetic data balancing

Beyond Rebalancing: Benchmarking Binary Classifiers Under Class Imbalance Without Rebalancing Techniques

Sep 09, 2025
AN
Ali Nawaz
🏛️ United Arab Emirates University | American University of the Middle East

In critical domains such as medical diagnosis, standard binary classifier evaluation under severe class imbalance often fails to reflect real-world robustness, especially when rebalancing techniques are inadmissible. Method: We propose a rebalancing-free robustness evaluation framework that synthesizes complex decision boundaries and adopts few-shot minority-class settings to emulate realistic extreme imbalance. We systematically benchmark TabPFN, ensemble boosting, one-class classification (OCC), and classical sampling methods across multiple real-world and synthetic datasets. Results: Traditional models exhibit significant performance degradation as minority-class prevalence decreases and data complexity increases; in contrast, TabPFN and ensemble methods demonstrate superior generalization and stability. This work is the first to reveal intrinsic robustness disparities among diverse models under unrebalanced conditions within a unified evaluation framework, establishing a new benchmark for imbalanced learning and offering actionable insights for practical deployment.

Assessing classifier robustness with reduced minority class sizesEvaluating binary classifiers without rebalancing under class imbalanceExploring performance across varying data complexities and imbalance scenarios

A Bilevel Optimization Framework for Imbalanced Data Classification

Oct 15, 2024
KM
Karen Medlin
🏛️ University of North Carolina at Chapel Hill | Argonne National Laboratory

Conventional resampling methods for class-imbalanced classification suffer from inherent limitations—oversampling introduces noise and boundary ambiguity, while undersampling discards informative majority-class samples, leading to information loss and underfitting. Method: This paper proposes an intelligent majority-class sample selection mechanism guided by model loss improvement. Its core innovation is a novel gradient-driven, differentiable bilevel optimization framework: the upper-level objective maximizes generalization performance, while the lower-level optimizes a differentiable loss improvement metric, enabling end-to-end, deterministic undersampling. Contribution/Results: By directly selecting discriminative majority-class instances—without synthesizing noisy minority samples—the method preserves data fidelity and decision boundary clarity. Evaluated on multiple benchmark datasets, it achieves up to a 10% absolute improvement in F1-score over state-of-the-art methods, significantly enhancing minority-class detection while maintaining majority-class accuracy.

Formulates bilevel optimization for optimal training subsetOptimizes majority data selection by improving model lossProposes undersampling method to avoid noise from synthetic data

Latest Papers

What's happening recently
View more

Concentration and excess risk bounds for imbalanced classification with synthetic oversampling

Oct 23, 2025
TA
Touqeer Ahmad
🏛️ Univ Rennes | Ensai | CNRS | CREST—UMR 9194 | Univ Angers | LAREMA | SFR MATHSTIC

This paper addresses the lack of theoretical foundations for synthetic oversampling methods such as SMOTE in imbalanced classification. We establish, for the first time, a statistical learning theory framework for such methods. By integrating uniform concentration inequalities with nonparametric estimation, we rigorously characterize the convergence between the empirical risk—computed over synthetically augmented data—and the population risk under the true data distribution, and derive a nonparametric excess risk bound for kernel classifiers. Our key contributions are: (1) the first unified concentration bound applicable to SMOTE-like oversampling schemes; (2) a theoretically grounded criterion for joint tuning of oversampling parameters (e.g., neighborhood size, synthesis ratio) and classifier hyperparameters; and (3) empirical validation demonstrating that the derived bound effectively predicts and guides generalization performance in practice.

Analyzing theoretical foundations of SMOTE for imbalanced classificationDeriving concentration bounds between synthetic and true minority distributionsProviding excess risk guarantees for kernel classifiers with synthetic data

Sampling Control for Imbalanced Calibration in Semi-Supervised Learning

Nov 24, 2025
ST
Senmao Tian
🏛️ Beijing Jiaotong University

In semi-supervised learning (SSL), distribution mismatch between labeled and unlabeled data exacerbates classification bias induced by class imbalance, yet existing methods conflate class imbalance with intrinsic learning difficulty and apply only coarse-grained corrections. This paper proposes SC-SSL, the first unified SSL framework that explicitly decouples these two factors. SC-SSL introduces a two-stage sampling control mechanism—adaptive resampling in the feature space and explicit classifier expansion in the logits space—complemented by bias-vector-driven logits calibration at inference time. This fine-grained balancing strategy jointly mitigates representation-level and prediction-level biases for minority classes. Extensive experiments across multiple benchmark datasets and diverse distribution shift settings demonstrate that SC-SSL consistently improves minority-class accuracy and overall class-balanced performance, outperforming state-of-the-art methods.

Addresses class imbalance in semi-supervised learning with distribution mismatchesCalibrates classifier logits by analyzing weight imbalance during inferenceMitigates feature-level imbalance for minority classes through adaptive sampling

To address classifier bias toward majority classes induced by class imbalance, this paper proposes a distribution-calibration-based synthetic sample generation method. The core innovation lies in estimating minority-class distribution parameters via a weighted Gaussian mixture model fitted on data from the proximity regions between majority and intermediate classes, while preserving semantic structure through an encoder-decoder network to generate high-fidelity synthetic samples; this strategy effectively mitigates overgeneralization of minority classes arising from sole reliance on majority-class modeling. Extensive experiments across multimodal datasets—including image, text, and tabular domains—demonstrate that the proposed method significantly outperforms mainstream baselines such as SMOTE, ADASYN, and CTGAN, achieving state-of-the-art performance in key metrics including F1-score and G-mean, with enhanced classification accuracy and robustness.

Addresses classifier bias from insufficient minority class dataGenerates synthetic samples through calibrated distribution parametersMitigates overgeneralization by leveraging neighboring class distributions

To address classification bias arising from the decoupling of learning optimization and model training in multi-class imbalanced classification, this paper proposes a density-aware and region-guided collaborative optimization Boosting framework. Methodologically, we introduce a novel noise-robust weight update mechanism that jointly incorporates density and confidence factors, and design a dynamic region partitioning strategy coupled with adaptive reweighted sampling—enabling end-to-end joint optimization of weight updates, region modeling, and sample selection. Technically, the framework integrates ensemble-based density estimation, confidence modeling, and dynamic sampling within a differentiable, trainable Boosting architecture. Extensive experiments on 20 public imbalanced datasets demonstrate significant improvements over eight state-of-the-art methods. The source code is publicly available.

Enhances model training with collaborative boosting approachIntegrates density and confidence factors for optimizationMitigates classification bias from class imbalance

The impact of class imbalance correction on model discriminative performance and probability calibration in clinical risk prediction remains unclear. This study systematically evaluates the effects of SMOTE, random oversampling (ROS), and random undersampling (RUS) across ten real-world clinical datasets using a range of linear and nonlinear models. Comprehensive comparisons are conducted using metrics including ROC-AUC, Brier score, and calibration intercept/slope. Results indicate that none of the three resampling methods significantly improve discrimination, yet all consistently degrade probability calibration—evidenced by increased Brier scores (0.029–0.080) and substantial shifts in calibration parameters—revealing systematic distortion in predicted risk estimates. These findings challenge the conventional use of resampling techniques in clinical prediction modeling.

class imbalanceclinical risk predictionmachine learning

Hot Scholars

FA

Faisal Ahmed

Samson Gemmell Chair of Child Health, University of Glasgow
Endocrinology
HW

Hassan Wasswa

University of New South Wales (UNSW)
Deep LearningInternet of ThingsCybersecurityComputer Vision
AJ

Alexis Joly

Research Director, Inria, Montpellier University, LIRMM
machine learningbiodiversityinformation retrievalplant identification
PB

Pierre Bonnet

Professor, Institut Pascal, clermont université
Electromagnetic compatibilityMaxwellnumerical methodsFVTD
AH

Ali Hamdi

Computer Science, MSA University
Computer VisionDeep LearningText Mining