🤖 AI Summary
This work addresses the challenge of cross-modal inconsistency in multimodal intent recognition, which often leads to semantic conflicts and information cancellation. To this end, we propose a cognition-inspired dual-pathway reasoning framework: an intuitive pathway establishes a consistent semantic foundation, while a reasoning pathway detects and mitigates high-level semantic conflicts, jointly modeling deep semantic associations. The framework integrates representation disentanglement, semantic prototype matching, statistical probability calibration, and a multi-view loss function, augmented with an inconsistency-aware mechanism and dynamic weight adjustment to effectively distinguish and fuse congruent and conflicting modalities. Experimental results demonstrate that our model achieves state-of-the-art performance on two benchmark datasets, significantly enhancing robustness against multimodal inconsistencies.
📝 Abstract
Multimodal Intent Recognition (MIR) aims to understand complex user intentions by leveraging text, video, and audio signals. However, existing approaches face two key challenges: (1) overlooking intricate cross-modal interactions for distinguishing consistent and inconsistent cues, and (2) ineffectively modeling multimodal conflicts, leading to semantic cancellation. To address these, we propose a novel Cognitive Dual-Pathway Reasoning (CDPR) framework, which constructs a stable semantic foundation via the intuition pathway and mitigates high-level semantic conflicts through the reasoning pathway, cooperatively establishing deep semantic relations. Specifically, we first employ a representation disentanglement strategy to extract modality-invariant and specific features. Subsequently, the intuition pathway aggregates cross-modal consensus using shared features for solid global representations. The reasoning pathway introduces an inconsistency perception mechanism, combining semantic prototype matching with statistical probability calibration to precisely quantify conflict severity, and dynamically adjusting the weights between both pathways. Furthermore, a multi-view loss function is adopted to alleviate modality laziness and learn structured features at different stages. Extensive experiments on two benchmarks show that CDPR achieves SOTA performance and superior robustness in mitigating multimodal inconsistency. The code is available at https://github.com/Hebust-NLP/CDPR.