Dynamic Alignment and Calibration for Multimodal Learning
This study addresses the limitations of static alignment in multimodal learning, which overlooks sample-level variations, and the failure of fusion weights to account for feature magnitude and confidence bias. To this end, we propose the ACML framework. It introduces a dynamic cross-modal triplet alignment strategy based on confidence gaps to achieve sample-level adaptive matching. Furthermore, it designs a discrepancy-aware attention calibration mechanism that jointly models feature magnitude and confidence discrepancies to optimize fusion weights. Extensive experiments on multiple benchmark datasets demonstrate that the proposed method significantly outperforms existing state-of-the-art approaches in both multimodal classification performance and robustness.