Learning What to Trust in Multimodal Learning under Noisy Supervision

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the underutilization of cross-modal information in noise detection for multimodal learning by proposing the REFINE framework. This method constructs a trustworthy representation space through the joint integration of fused and unimodal representations to purify supervision signals. Furthermore, it provides the first theoretical analysis revealing the intrinsic relationship between representation structure and noise detection capability, dynamically selecting the optimal trustworthy subspace for each class based on discriminative eigenvectors. Experimental results demonstrate that the proposed framework significantly outperforms existing baselines across multiple tasks, effectively mitigating label noise interference and enhancing model generalization performance.
📝 Abstract
Multimodal classification processes and relates information from multiple modalities to achieve more accurate predictions. However, existing methods typically rely on high-quality ground-truth labels, which are difficult to obtain in real-world scenarios. While sample-selection methods for learning with noisy labels aim to identify correctly labeled examples from noisy data, traditional methods primarily focus on unimodal settings and fail to exploit multimodal information fully. This motivates us to build a more reliable noise detector in multimodal learning. To this end, we theoretically analyze the relationship between representation structure and noise detection capability. Based on this analysis, we propose REFINE, which is a multimodal label-noise detection framework that jointly uses fused and unimodal representations for label-noise detection. Specifically, REFINE constructs discriminative eigenvectors through discriminative analysis of the target and background classes and selects trusted representation spaces with better noise detection capability for each class. Within each trusted space, REFINE measures the alignment between each instance representation and the discriminative eigenvectors. It then combines the subsets selected from these spaces. The combined set provides cleaner supervision for updating the multimodal classifier, thereby reducing the influence of mislabeled examples during training and improving model generalization. Extensive experiments across diverse tasks demonstrate REFINE's superiority compared to baseline methods. The source code will be publicly available.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Learning
Noisy Labels
Sample Selection
Noise Detection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Learning
Noisy Labels
Sample Selection
Discriminative Analysis
Representation Alignment
💼 Related Jobs
No related jobs found.