🤖 AI Summary
This study addresses the performance degradation in multimodal sentiment analysis caused by noise interference and missing modalities. To overcome the limitations of conventional approaches that handle these issues in isolation, this work proposes a Reliability-aware Cross-sample Enhancement (RCE) framework, introducing a pioneering joint mitigation strategy. Specifically, RCE employs an adaptive variational information bottleneck to suppress noise and leverages high-confidence neighbor retrieval for cross-sample enhancement to compensate for missing information. Furthermore, a multi-level reliability-aware fusion mechanism is designed to ensure robust integration. Experimental results demonstrate that RCE significantly outperforms existing state-of-the-art methods across complete, noisy, and modality-missing scenarios, validating its effectiveness and robustness in challenging multimodal settings.
📝 Abstract
Multimodal Sentiment Analysis (MSA) aims to infer human emotions from multiple modalities such as text, audio, and vision. In practice, inputs are often corrupted by noise and missing modalities, which degrades performance. Existing methods typically address these challenges in isolation, limiting their effectiveness in realistic settings. To address this limitation, we propose a Reliability-aware Cross-sample Enhancement (RCE) framework. Specifically, RCE first introduces an adaptive variational information bottleneck to model modality-wise uncertainty and perform quality-aware information compression, thereby suppressing redundant noise in unreliable modalities. Furthermore, we design a reliability-aware cross-sample enhancement strategy that retrieves high-confidence, semantically consistent neighbors from a large candidate pool to enrich and calibrate current representations, effectively alleviating information deficiency caused by missing modalities. Building upon this, RCE integrates cross-modal interactions with a multilevel reliability-aware fusion mechanism to adaptively aggregate information across modalities and enhancement stages, leading to more robust multimodal representations. Extensive experiments demonstrate that RCE consistently outperforms state-of-the-art methods across full, noisy, and missing-modality settings.