Bringing Clustering to MLL: Weakly-Supervised Clustering for Partial Multi-Label Learning

📅 2026-04-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the noise inherent in partial multi-label learning, where candidate label sets are contaminated with both relevant and irrelevant labels. To tackle this challenge, the authors propose a novel weakly supervised clustering approach that uniquely decomposes the cluster membership matrix into two components: a normalized Π component and an F component that preserves the binary nature of multi-label assignments. This decomposition enables the first effective integration of clustering with multi-label learning. The method employs a three-stage pipeline—prototype learning, confidence-adaptive weak supervision construction, and iterative clustering refinement—to achieve robustness against label noise. Extensive experiments on 24 benchmark datasets demonstrate that the proposed approach significantly outperforms six state-of-the-art methods across all evaluation metrics.

Technology Category

Machine Learning: Multi-class/Multi-label Learning & Extreme ClassificationComputer Vision: Multi-modal VisionMultiagent Systems: Multiagent Learning

Application Category

Web Mining and Content Analysis: Normalization, clustering, classification, and summarization of Web textGraph Algorithms and Modeling for the Web: Algorithms and analysis for incomplete, noisy, or partially observed Web-related graphsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
📝 Abstract
Label noise in multi-label learning (MLL) poses significant challenges for model training, particularly in partial multi-label learning (PML) where candidate labels contain both relevant and irrelevant labels. While clustering offers a natural approach to exploit data structure for noise identification, traditional clustering methods cannot be directly applied to multi-label scenarios due to a fundamental incompatibility: clustering produces membership values that sum to one per instance, whereas multi-label assignments require binary values that can sum to any number. We propose a novel weakly-supervised clustering approach for PML (WSC-PML) that bridges clustering and multi-label learning through membership matrix decomposition. Our key innovation decomposes the clustering membership matrix $\mathbf{A}$ into two components: $\mathbf{A} = \mathbf{\Pi} \odot \mathbf{F}$, where $\mathbf{\Pi}$ maintains clustering constraints while $\mathbf{F}$ preserves multi-label characteristics. This decomposition enables seamless integration of unsupervised clustering with multi-label supervision for effective label noise handling. WSC-PML employs a three-stage process: initial prototype learning from noisy labels, adaptive confidence-based weak supervision construction, and joint optimization via iterative clustering refinement. Extensive experiments on 24 datasets demonstrate that our approach outperforms six state-of-the-art methods across all evaluation metrics.
Problem

Research questions and friction points this paper is trying to address.

partial multi-label learning
label noise
clustering
multi-label learning
weakly-supervised learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

weakly-supervised clustering
partial multi-label learning
membership matrix decomposition
label noise
multi-label learning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yu Chen
School of Automation, Guangdong University of Technology
W
Weijun Lv
School of Automation, Guangdong University of Technology
Y
Yue Huang
School of Automation, Guangdong University of Technology
X
Xuhuan Zhu
School of Automation, Guangdong University of Technology
F
Fang Li
School of Automation, Guangdong University of Technology