multi-label classification

Design and build models, training objectives, and data-processing pipelines that predict multiple non-mutually-exclusive labels for each instance, including architectures and inference methods that capture label correlations and structured label outputs. Create and analyze multi-label loss functions, evaluation procedures (e.g., micro-F1 and other multi-label metrics), and practical adaptations of base learners or recommender components to produce accurate, calibrated multi-label predictions.

multi-labelclassification

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.44
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$227K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Multi-Label Contrastive Learning : A Comprehensive Study

Nov 27, 2024
AA
Alexandre Audibert
🏛️ Université Grenoble Alpes

This work systematically investigates the effectiveness and limitations of multi-label contrastive learning across diverse settings. Addressing challenges in multi-label classification—particularly difficulty in modeling label dependencies and poor robustness under few-shot conditions—we propose a supervised contrastive loss specifically designed for multi-label scenarios. We theoretically and empirically establish that its performance gains stem from the synergistic interplay between explicit label interaction modeling and gradient-robust optimization. Extensive cross-modal experiments on computer vision and natural language processing benchmarks demonstrate consistent improvements in Macro-F1 (+2.3% on average) under large-scale label spaces (>100 classes) and moderate data regimes, while effectively capturing semantic label correlations. Crucially, we delineate its applicability boundaries: marginal gains are observed for extremely small label sets (<10 classes) or ranking-oriented metrics (e.g., Recall@k). Our findings provide both theoretical insights and practical guidelines for deploying contrastive learning in multi-label settings.

Adaptability and PerformanceContrastive LearningMultilabel Learning

Rethinking Consistent Multi-Label Classification under Inexact Supervision

Oct 05, 2025
WW
Wei Wang
🏛️ The University of Tokyo | RIKEN Center for Advanced Intelligence Project

Existing methods for partial multi-label learning (PML) and complementary multi-label learning (CML)—two prominent weakly supervised multi-label classification paradigms—rely either on precise modeling of the label generation process or strong uniformity assumptions about label distributions, both of which are frequently violated in practice. Method: We propose the first unified, unbiased, and statistically consistent learning framework that requires neither label-generation modeling nor uniformity assumptions, enabling joint unbiased risk estimation for both PML and CML. Our approach constructs risk estimators via first- and second-order moment correction. Contribution/Results: We rigorously establish statistical consistency and convergence rates of the proposed estimators under standard evaluation metrics—including Hamming loss, Jaccard index, and F1 score. Extensive experiments on multiple benchmark datasets demonstrate significant improvements over state-of-the-art methods, validating both effectiveness and robustness.

Addresses weakly supervised multi-label classification with inexact supervisionHandles both partial and complementary multi-label learning problemsProposes consistent approaches without relying on label generation estimation

Constraint-aware Learning of Probabilistic Sequential Models for Multi-Label Classification

Jul 20, 2025
MB
Mykhailo Buleshnyi
🏛️ Ukrainian Catholic University | HUN-REN Alfréd Rényi Institute of Mathematics | Eötvös Loránd University | University of Oxford

To address the challenge of modeling and enforcing logical constraints in multi-label classification with large label sets, this paper proposes a joint distribution framework integrating single-label classifiers with probabilistic sequence models. Methodologically, logical constraints are explicitly encoded into both the training objective—via a constraint-guided loss function—and the inference procedure—through constraint-enforced decoding—while sequence modeling captures high-order label dependencies. The key contribution is the first end-to-end framework that simultaneously incorporates constraint information into both training and inference stages, ensuring logical consistency throughout the pipeline. Experiments on multiple large-scale multi-label benchmarks demonstrate significant improvements: average classification accuracy increases by 2.1%, and constraint satisfaction rates improve by up to 37.5%, all while maintaining efficient inference.

Enforcing constraints during both training and inferenceHandling large sets of labels with logical constraintsModeling label correlations via sequential probabilistic models

This study addresses the challenge of capturing high-order label dependencies in multi-label classification by proposing the HyperLabel framework. The method constructs a label hypergraph with samples as hyperedges, providing structural priors that transcend pairwise interactions. Building upon hypergraph neural networks, an encoder-decoder architecture is designed to achieve unified cross-modal fusion of feature and structural information through cross-attention and bidirectional message passing mechanisms. Experimental results demonstrate that the proposed framework achieves state-of-the-art performance across seven benchmark datasets. Notably, it yields substantial improvements of 10.3% and 8.2% in Macro-F1 on the Delicious and Bibtex datasets, respectively. These findings validate the effectiveness of explicitly modeling complex label co-occurrence patterns for advancing multi-label classification performance.

High-order dependenciesHypergraphLabel correlation

Latest Papers

What's happening recently
View more

This work investigates whether the reported performance gains of existing multi-label node classification methods stem from specialized designs or merely from insufficient optimization of classical baselines. To address this, the authors systematically enhance general-purpose GNN architectures—such as GCN, SSGConv, and GCNII—by integrating standard techniques including normalization, Dropout, and residual connections to construct strong baselines. Extensive experiments demonstrate that these well-tuned classical models outperform current specialized approaches on four out of five benchmark datasets and achieve state-of-the-art results across various experimental settings. These findings underscore the critical importance of employing rigorously optimized baselines in multi-label graph learning research to ensure meaningful methodological comparisons.

graph neural networkslabel dependenciesmulti-label node classification

Existing LLM-driven feature engineering methods are not designed for multi-label learning, thus failing to model label dependencies and lacking task-specificity. To address this, we propose FEAML—a novel framework that pioneers the integration of LLM-based code generation into multi-label settings. FEAML automatically constructs highly discriminative features by jointly leveraging metadata and label co-occurrence matrices. It introduces label-dependency-aware prompt engineering and a Pearson correlation-based redundancy detection mechanism, coupled with closed-loop optimization guided by classification accuracy. This yields an interpretable, low-redundancy, and self-optimizing feature generation paradigm. Extensive experiments on multiple standard multi-label benchmark datasets demonstrate that FEAML significantly outperforms conventional feature engineering approaches, achieving substantial average improvements in classification accuracy—thereby validating its effectiveness and generalizability.

FEAML automates feature engineering for multi-label classification tasks.It models label dependencies using metadata and co-occurrence matrices.The method integrates feedback to optimize LLM-generated features iteratively.

This work addresses the dual incompleteness problem in multi-view multi-label learning, where both views and labels may be simultaneously missing—a scenario in which existing methods struggle to learn stable and discriminative shared representations due to the lack of explicit structural constraints. To tackle this challenge, the authors propose a structured consistent representation learning framework that learns discrete consensus representations through a shared multi-view codebook and cross-view reconstruction. At the decision level, a label-correlation-aware view-weighted fusion mechanism is introduced to leverage structural dependencies among labels. Furthermore, a teacher-guided self-distillation architecture is designed to distill global knowledge back into individual view-specific branches, thereby enhancing generalization. Extensive experiments on five benchmark datasets demonstrate that the proposed method significantly outperforms state-of-the-art approaches, confirming its effectiveness and robustness under dual missingness conditions.

consistent representationdual-missing scenarioincomplete multi-view

This work addresses the limitations of traditional gradient boosting methods—such as limited expressiveness, sensitivity to hyperparameters, and difficulty in continual learning—when handling both structured and unstructured data. The authors propose Multiplicative Additive Neural Networks (MANN), which replace decision trees with shallow neural networks as base learners to enable unified modeling of images, audio, and tabular data. Notably, MANN incorporates capsule networks for the first time into feature extraction for structured data. Embedded within an enhanced gradient boosting framework, MANN integrates a continual learning mechanism and regularization strategies, substantially reducing sensitivity to learning rate and iteration count. Experimental results demonstrate that MANN outperforms strong baselines such as XGBoost across multiple benchmark datasets, exhibiting superior generalization and multimodal adaptability.

gradient boostinghyperparameter sensitivityoverfitting

This work addresses the challenge of multi-label classification, where complex inter-label dependencies and highly nonlinear relationships with features hinder effective modeling by existing methods. The authors propose the first extension of Bayesian Additive Regression Trees (BART) to multi-label classification by introducing latent continuous variables that are thresholded to produce discrete labels, while explicitly capturing label correlations through a multivariate normal distribution. This unified framework jointly models nonlinear effects, label dependencies, and predictive uncertainty, with inference performed via Markov chain Monte Carlo (MCMC). Experimental results demonstrate that the proposed model significantly outperforms current approaches on simulated data, achieving prediction accuracy nearly matching that of an oracle model, and provides well-calibrated conditional probabilities for label combinations along with reliable uncertainty quantification.

Bayesian Additive Regression TreesLabel CorrelationMCMC

Hot Scholars

MS

Masashi Sugiyama

Director, RIKEN Center for Advanced Intelligence Project / Professor, The University of Tokyo
Machine LearningData MiningArtificial Intelligence
MK

Ming-Kun Xie

RIKEN Center for Advanced Intelligence Project
machine learning
JW

Jie Wen

Associate Professor, North University of China(NUC)
Quantum ControlPrognostic and Health Management
AJ

Alexis Joly

Research Director, Inria, Montpellier University, LIRMM
machine learningbiodiversityinformation retrievalplant identification
ZK

Zhiqiang Kou

Ph.D. Student at Southeast University, Internship at RIKEN AIP
Machine learning