class-weighted training

Designs and implements training and fine‑tuning procedures that incorporate class- or example-specific weights or costs into loss functions and boosting algorithms (e.g., weighted cross-entropy, class‑balanced terms, sample weights in XGBoost, or post‑training weighted binary loss). Builds and evaluates models and training pipelines to mitigate label imbalance and prioritize minority or high‑importance classes, measuring effects on balanced metrics such as macro F1 or balanced accuracy.

class-weightedtraining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.21
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the unreliability of conclusions regarding class-imbalance methods derived from single-dataset evaluations. Employing a leakage-free nested cross-validation protocol across 45 binary classification tasks, we conducted large-scale experiments to reassess these techniques. Results reveal that threshold tuning benefits exhibit non-monotonic variation with imbalance ratios and refute the hypothesis that calibration error predicts tuning gains. Furthermore, Random Forest combined with SMOTE demonstrated significant effectiveness across multiple tasks, highlighting the limitations of findings based solely on single fraud datasets. This work systematically clarifies the true utility and applicability boundaries of resampling and threshold tuning, providing robust empirical evidence for the reliable evaluation of imbalanced classification methods.

GeneralizabilityImbalanced ClassificationResampling

FILM: Framework for Imbalanced Learning Machines based on a new unbiased performance measure and a new ensemble-based technique

Mar 06, 2025
AG
Antonio Guillén-Teruel
🏛️ Universidad de Murcia | University College London

To address the substantial bias in standard evaluation metrics (e.g., accuracy, F1-score) and consequent unreliability in model selection under binary class imbalance, this paper proposes the Unbiased Imbalance-aware Criterion (UIC) and the Imbalance-Penalized Integrated Prediction (IPIP) ensemble method, unified within the FILM framework. UIC introduces imbalance-sensitive penalization and multi-metric weighted aggregation—yielding theoretically grounded, statistically significant reduction in minority-class bias (p < 10⁻⁴). IPIP mitigates distribution shift via consistent data partitioning and ensemble integration of base learners (random forests and logistic regression). Empirical evaluation across seven real-world imbalanced datasets demonstrates that IPIP achieves significantly higher UIC scores than state-of-the-art imbalance learning methods on three datasets. The FILM framework is publicly released as an open-source R package.

Addresses bias in evaluation metrics for imbalanced datasets.Introduces IPIP algorithm to improve classification in imbalanced data.Proposes Unbiased Integration Coefficients (UIC) for consistent model evaluation.

This study addresses the issue of model bias toward majority classes in imbalanced classification by formally framing it as a label shift domain adaptation problem between the source distribution (observed data) and the target distribution (balanced evaluation distribution). The authors introduce the concept of “transfer cost” and provide theoretical analysis showing that SMOTE incurs higher transfer cost than random oversampling methods such as Bootstrap in medium- to high-dimensional spaces. Building on this framework, they integrate minority class distribution estimation into data augmentation and empirically demonstrate that random oversampling generally outperforms SMOTE in such settings. These findings offer both theoretical justification and practical guidance for selecting oversampling strategies in imbalanced classification tasks.

classification imbalancelabel shiftoversampling

A Unified Generalization Analysis of Re-Weighting and Logit-Adjustment for Imbalanced Learning

Oct 07, 2023
ZW
Zitai Wang
🏛️ State Key Laboratory of AI Safety | Institute of Computing Technology, Chinese Academy of Sciences | Peng Cheng Laboratory | University of Chinese Academy of Sciences | School of Computer Science and Technology | School of Cyber Science and Technology | Shenzhen Campus of Sun Yat-sen University | Key Laboratory of Big Data Mining and Knowledge Management (BDKM)

Empirical Risk Minimization (ERM) suffers from degraded generalization under long-tailed class distributions. Method: This paper proposes a data-dependent shrinkage technique and establishes the first fine-grained, class-aware unified generalization upper bound. Unlike conventional coarse-grained analyses relying on global statistics, our bound explicitly quantifies how class-specific terms influence generalization error. Contribution/Results: The bound provides the first systematic theoretical explanation of the intrinsic mechanisms underlying reweighting and logit adjustment—resolving several counterintuitive empirical observations. Leveraging this theory, we design a principled learning algorithm that significantly improves minority-class accuracy on standard long-tailed benchmarks—including CIFAR-10-LT and ImageNet-LT—outperforming state-of-the-art methods.

Addresses class imbalance in datasets affecting ERM generalization.Analyzes localized properties for loss-oriented imbalanced learning methods.Develops a unified perspective and algorithm for improved model performance.

Restoring balance: principled under/oversampling of data for optimal classification

May 15, 2024
EL
Emanuele Loffredo
🏛️ PSL University | Sorbonne University | Université Paris-Cité

Linear classifiers (e.g., SVM) suffer from degraded generalization performance on high-dimensional imbalanced data. Method: We establish a high-dimensional asymptotic theoretical framework and, for the first time, rigorously derive analytical expressions for the generalization error under undersampling and oversampling. Our approach integrates random matrix theory, high-dimensional statistical learning, and unsupervised probabilistic modeling–driven resampling. Contribution/Results: We quantify how resampling efficacy depends on the first- and second-order statistics of the data and the choice of evaluation metric. Crucially, we prove—and empirically verify—that hybrid sampling consistently outperforms either undersampling or oversampling alone. Extensive numerical experiments and evaluations on real-world datasets—including deep neural network features—demonstrate strong agreement between theoretical predictions and empirical results, with substantial improvements in minority-class classification accuracy. This work provides an interpretable, generalizable, and principle-based foundation for data rebalancing in high dimensions.

High-dimensional dataImbalanced datasetsLinear classifiers

Latest Papers

What's happening recently
View more

To address classification bias arising from the decoupling of learning optimization and model training in multi-class imbalanced classification, this paper proposes a density-aware and region-guided collaborative optimization Boosting framework. Methodologically, we introduce a novel noise-robust weight update mechanism that jointly incorporates density and confidence factors, and design a dynamic region partitioning strategy coupled with adaptive reweighted sampling—enabling end-to-end joint optimization of weight updates, region modeling, and sample selection. Technically, the framework integrates ensemble-based density estimation, confidence modeling, and dynamic sampling within a differentiable, trainable Boosting architecture. Extensive experiments on 20 public imbalanced datasets demonstrate significant improvements over eight state-of-the-art methods. The source code is publicly available.

Enhances model training with collaborative boosting approachIntegrates density and confidence factors for optimizationMitigates classification bias from class imbalance

This study addresses the critical challenge of missed diagnosis of high-risk cardiac discharge phenotypes in real-world clinical settings, where data scarcity and class imbalance are prevalent. To this end, the authors propose a clinically risk-aligned, class-weighted XGBoost framework that integrates instance-level class weighting guided by clinical priorities, explicit modeling of missing values via missingness indicators, and a class-level error auditing mechanism. This approach enhances the model’s ability to identify minority high-risk phenotypes while preserving interpretability. Evaluated under five-fold stratified cross-validation, the proposed method significantly outperforms state-of-the-art tree-based models, ensemble techniques, and neural network baselines across multiple metrics, including Accuracy, Macro-F1, Balanced Accuracy, and Prioritized F1.

cardiac phenotypingclass imbalancehigh-risk phenotype

This work addresses the instability and performance degradation commonly observed during fine-tuning of pre-trained models, which often stems from gradient cancellation leading to optimization collapse. To mitigate this issue, the paper introduces, for the first time in the context of fine-tuning, a dynamic gradient scaling mechanism, proposing the Dynamic Scaled Gradient Descent (DSGD) algorithm. DSGD adaptively attenuates the gradient magnitudes of correctly classified samples, thereby effectively alleviating gradient cancellation. The method substantially enhances fine-tuning stability and robustness, consistently reducing performance variance and achieving higher accuracy than existing approaches across multiple benchmark datasets and large-scale models.

class imbalancefine-tuninggradient collapse

This work addresses the tendency of neural networks to bias toward majority classes in imbalanced datasets by proposing a novel loss function that, for the first time, incorporates cardinality-based invariants from metric space—such as magnitude and spread—into the training process to enhance effective data diversity. By explicitly quantifying and optimizing the geometric structure of sample distributions, the method significantly improves the model’s ability to recognize minority classes. Experiments on both synthetic and real-world materials science imbalanced datasets demonstrate consistent and substantial gains in both minority-class performance and overall evaluation metrics, thereby validating the effectiveness and generalizability of the proposed approach.

class imbalanceclassifier performanceimbalanced datasets

Hot Scholars

HI

Hitoshi Iyatomi

Professor, Hosei University, Japan
deep learningcomputer visionmachine learningmedical engineering
YL

Yong Li

Institue of Software, Chinese Academy of Sciences
Automata theoryModel checking
DZ

Dan Zeng

Sun Yat-sen University
Biometricscomputer visiondeep learning
SG

Shiming Ge

Institute of Information Engineering, Chinese Academy of Sciences
Computer VisionArtificial Intelligence
ME

Md. Ehsanul Haque

East West University
Machine Learning in CybersecurityMachine LearningImage ProcessingHealth Informatics