A Comprehensive Survey on Imbalanced Data Learning

📅 2025-02-13
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Imbalanced data are pervasive in multimodal real-world scenarios, severely undermining the generalizability and fairness of machine learning models. To address this, this work introduces the first systematic methodology taxonomy for imbalanced learning across heterogeneous data modalities—specifically image, text, and time-series data—categorized into four unified paradigms: data rebalancing, feature representation learning, training strategy design, and ensemble learning. Through comprehensive survey analysis, cross-modal modeling, meta-analysis, and empirical evaluation of open-source tools, we identify the applicability boundaries and shared challenges of each paradigm across data formats. We propose a structured analytical framework to advance standardized understanding and highlight critical gaps in the open-source ecosystem. This paper constitutes the first holistic, format-agnostic survey and methodological integration for imbalanced learning, establishing foundational principles for future research and practice.

Technology Category

Machine Learning: Multimodal LearningComputer Vision: Multi-modal VisionData Mining & Knowledge Management: Mining of Visual, Multimedia & Multimodal Data

Application Category

Web Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web dataUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
📝 Abstract
With the expansion of data availability, machine learning (ML) has achieved remarkable breakthroughs in both academia and industry. However, imbalanced data distributions are prevalent in various types of raw data and severely hinder the performance of ML by biasing the decision-making processes. To deepen the understanding of imbalanced data and facilitate the related research and applications, this survey systematically analyzing various real-world data formats and concludes existing researches for different data formats into four distinct categories: data re-balancing, feature representation, training strategy, and ensemble learning. This structured analysis help researchers comprehensively understand the pervasive nature of imbalance across diverse data format, thereby paving a clearer path toward achieving specific research goals. we provide an overview of relevant open-source libraries, spotlight current challenges, and offer novel insights aimed at fostering future advancements in this critical area of study.
Problem

Research questions and friction points this paper is trying to address.

Addressing imbalanced data distributions
Improving ML performance bias
Systematizing imbalanced data learning methods
Innovation

Methods, ideas, or system contributions that make the work stand out.

Data re-balancing techniques
Feature representation methods
Ensemble learning strategies
🔎 Similar Papers