Dynamic Alignment and Calibration for Multimodal Learning

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of static alignment in multimodal learning, which overlooks sample-level variations, and the failure of fusion weights to account for feature magnitude and confidence bias. To this end, we propose the ACML framework. It introduces a dynamic cross-modal triplet alignment strategy based on confidence gaps to achieve sample-level adaptive matching. Furthermore, it designs a discrepancy-aware attention calibration mechanism that jointly models feature magnitude and confidence discrepancies to optimize fusion weights. Extensive experiments on multiple benchmark datasets demonstrate that the proposed method significantly outperforms existing state-of-the-art approaches in both multimodal classification performance and robustness.
📝 Abstract
Dynamic multimodal learning aims to learn robust representations by adaptively modeling information discrepancies across modalities. However, existing methods still suffer from two limitations: (i) static cross-modal alignment strategies usually impose uniform constraints on all samples while overlooking sample-wise variations, potentially leading to unreasonable over-alignment; and (ii) confidence- or uncertainty-aware fusion methods often fail to adequately account for feature magnitude and confidence differences across modalities. For modality pairs with significant feature magnitude differences or small confidence gaps, it might be unreliable to strictly align fusion weights according to confidence. To address these issues, we propose an Alignment- and Calibration-driven Multimodal Learning framework (ACML). Specifically, ACML incorporates a dynamic cross-modal triplet alignment module, which enforces strong semantic consistency for high-confidence positive pairs while encouraging diverse representation learning between high- and low-confidence positive pairs according to their confidence gaps. Additionally, ACML introduces a difference-aware attention calibration strategy that adaptively adjusts attention regularization based on feature magnitude and confidence differences across modalities, thereby mitigating biases caused by unreasonable fusion constraints. Extensive experiments on multiple multimodal benchmark datasets demonstrate that ACML consistently achieves superior performance and robustness over recent state-of-the-art methods.
Problem

Research questions and friction points this paper is trying to address.

multimodal learning
cross-modal alignment
confidence calibration
feature magnitude
fusion
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dynamic Multimodal Learning
Cross-modal Triplet Alignment
Attention Calibration
Difference-aware Fusion
Multimodal Representation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jinghao Xu
School of Computer Science and Engineering, University of Electronic Science and Technology of China, Chengdu, China
Zhenhua Guo
Zhenhua Guo
Tianyijiaotong Technology Ltd., China
X
Xiaofeng Zhu
School of Computer Science and Technology, Hainan University, Haikou, China
Xiaoshuang Shi
Xiaoshuang Shi
University of Electronic Science and Technology of China
Machine learningcomputer visionmedical image analysis