Score
Designs and implements attention-based modules that fuse information across multiple modalities by computing cross-attention and cross-encoder interactions between modality-specific token or feature streams, supporting per-layer fusion, multi-token cross-attention, cross-source/global-context aggregation, and mirror/encoder variants. Also builds and analyzes modality-aware weighting and aggregation strategies (learnable modality weights, blur-aware or bi-temporal cross-attention), plus compact attention-driven classifier heads or layer-fusion strategies to produce joint multimodal representations or decisions.
This paper systematically analyzes four core challenges in multimodal alignment and fusion—cross-modal misalignment, semantic modality gap, computational bottlenecks, and data noise/heterogeneity—that hinder model robustness, generalizability, and scalability. To address them, the authors propose a unified taxonomy covering 200+ studies, revealing an integrated “alignment–fusion” co-design paradigm. They further introduce three principled solution pathways: (1) noise-robust learning, (2) heterogeneous representation modeling, and (3) few-shot cross-modal transfer—unifying contrastive learning, cross-modal attention, latent-space alignment, graph neural networks, and differentiable architecture search. Extensive experiments on social media analysis, medical imaging, and sentiment recognition validate the efficacy of the framework. The work distills actionable design principles for scalable, robust, and generalizable multimodal learning, establishing a new benchmark for both theoretical research and industrial deployment.
Existing research lacks a systematic analysis of modality–language backbone integration mechanisms in multimodal large language models (MLLMs). Method: We systematically survey 125 MLLMs published between 2021 and 2025, proposing the first large-language-model-centric three-dimensional taxonomy—architectural integration, representation learning, and training paradigms—to unify cross-modal alignment pathways. Leveraging bibliometric analysis, architectural decoupling, and fine-grained modeling of embeddings and loss functions, we characterize evolutionary patterns in fusion granularity, joint representation design, and objective function development. Contribution/Results: This work fills a critical theoretical gap in structured analysis of MLLM fusion mechanisms, identifies shared bottlenecks—including misaligned modality-specific feature hierarchies and suboptimal cross-modal optimization objectives—and delivers a reusable theoretical framework and practical guidelines for developing robust, scalable multimodal foundation models.
Contemporary multimodal large language models (MLLMs) suffer from inconsistent cross-modal attention and progressive layer-wise attenuation, hindering fine-grained perception, cognition, and affective understanding in advanced multimodal tasks. To address these limitations, we propose Modular Dual-path Attention (MODA), a novel architecture featuring three key innovations: (1) a “correct-post-alignment” strategy that decouples modality alignment from cross-layer token mixing; (2) adaptive masked attention enabling modality-specific interaction patterns; and (3) a unified bimodal embedding space constructed via foundational vector mapping. MODA preserves semantic fidelity while enhancing cross-modal coherence across layers. We comprehensively evaluate MODA on 21 diverse multimodal benchmarks—including visual reasoning, emotion recognition, and compositional understanding—demonstrating consistent and significant improvements over state-of-the-art MLLMs. All code and interactive demos are publicly released.
Multimodal learning faces two key challenges: intra-modal noise corrupting joint representations and loss of modality-specific discriminative information during fusion. To address these, we propose a multi-task multimodal semantic understanding framework. Our method employs (1) a cross-modal relational graph that implicitly models inter-modal dependencies—thereby avoiding explicit interaction and mitigating noise propagation—and (2) a Hierarchical Interactive Modality-wise Attention (HIMA) mechanism that enhances intra-modal feature extraction and discriminative modeling prior to late fusion. By integrating graph neural networks with attention mechanisms, the framework enables neighborhood-driven feature reconstruction and fine-grained semantic modeling. Extensive experiments on three public benchmarks demonstrate substantial improvements in both accuracy and robustness over state-of-the-art methods, validating the effectiveness of our approach in preserving modality-specific semantics while enabling robust cross-modal integration.
To address insufficient cross-modal interaction and limited fine-grained classification accuracy in multimodal sentiment analysis, this paper proposes a dynamic sentiment analysis model based on early feature fusion and a lightweight cross-modal attention mechanism. Methodologically, it adopts a Transformer architecture with a novel cross-modal interaction module tailored for multi-head attention, enabling joint encoding of textual, acoustic, and visual features at the input layer. It is the first work to systematically validate the significant superiority of early fusion on the CMU-MOSEI benchmark. Experiments show that the model achieves 72.39% accuracy—over 4 percentage points higher than representative late-fusion baselines—while reducing parameter count by 18%, thus balancing performance and efficiency. Key contributions are: (1) empirical confirmation of early fusion’s positive impact on multimodal sentiment representation; and (2) a low-overhead, highly compatible cross-modal attention strategy.
Existing audio-visual fusion methods struggle to balance cross-modal dependency modeling with computational efficiency, limiting the scalability of multi-scale architectures. To address this challenge, this work proposes SNNergy, a novel framework that achieves hierarchical multi-scale cross-modal fusion with linear complexity for the first time. At its core lies the CMQKA mechanism, which leverages event-driven binary spiking operations to construct an efficient bidirectional Query-Key attention and integrates a learnable residual fusion strategy. Evaluated on benchmark datasets including CREMA-D, AVE, and UrbanSound8K-AV, the proposed method attains state-of-the-art performance, significantly outperforming existing approaches while demonstrating exceptional energy efficiency.
RGB-D multimodal fusion mechanisms have long suffered from poor interpretability, and the fundamental nature of cross-modal complementarity remains unclear. Method: This paper establishes the first interpretability analysis framework specifically targeting the fusion process, introducing a joint metric of semantic variance and feature similarity to systematically characterize cross-modal representation consistency, intra-modal evolutionary patterns, and collaborative optimization logic. Through cross-layer feature comparison and quantitative semantic analysis, we identify a prevalent imbalance between consistency and specificity in mainstream fusion strategies. Contribution/Results: We formalize a “specificity-driven inference under consistency constraints” principle that explicates cross-modal complementarity. Our framework provides both theoretical foundations and a verifiable evaluation paradigm for designing trustworthy, generalizable multimodal fusion models.
This work addresses the challenge that certain modalities may introduce interference under specific inputs in multimodal fusion, thereby degrading model performance. To mitigate this issue, the authors propose a plug-and-play pre-fusion calibration module that leverages cross-modal summary contrast to extract supportive and conflicting cues, generating instance-level and dimension-level modulation signals. These signals dynamically enhance beneficial features while suppressing misleading information. The method achieves, for the first time, fine-grained, conflict-aware modulation prior to fusion and is compatible with both sequential and convolutional architectures. It consistently improves performance across five benchmark tasks—including emotion understanding and action recognition—and demonstrates enhanced robustness and consistency under modality missingness and data perturbations.
Existing multimodal fusion approaches often suffer from modality overfitting or excessive specialization, which compromises the generalization capability of cross-modal object detection. To address this, this work proposes an attention-driven complementary resampling framework that enhances cross-modal feature interaction through a shared channel-spatial attention mechanism. The method blurs modality boundaries via a semantic mask exchange strategy and introduces a learnable channel competition mechanism to enable dynamic, channel-wise sampling and aggregation. Notably, the framework operates without requiring explicit modality labels, thereby effectively promoting complementarity and generalization across modalities. Experimental results demonstrate that the proposed model consistently outperforms state-of-the-art methods on multiple benchmark datasets, confirming its efficacy and superiority.
This work systematically investigates the core mechanisms of modality interaction in multimodal pretraining, focusing on knowledge transfer, synergistic effects, and fusion timing. Through experiments on both synthetic and large-scale real-world datasets, it provides the first empirical evidence of asymmetric cross-modal knowledge flow and demonstrates that data complexity governs whether modalities exhibit synergy or competition. The study further validates that early unified fusion consistently outperforms late alignment. Leveraging an architecture featuring shared attention and normalization layers with modality-specific feedforward components, the proposed approach is evaluated on a 13.5B mixture-of-experts model trained on 2 trillion tokens, confirming its effectiveness. Additionally, the paper introduces a highly efficient pretraining strategy that achieves strong generative performance using only 5% of the typical computational budget.
Existing Transformer-based multimodal models struggle to efficiently capture high-order modality interactions, often being limited to pairwise interactions or suffering from quadratic computational complexity with respect to the number of modalities. This work proposes GRAMformer, a novel multimodal Transformer architecture centered on a Volume-based Multimodal Attention (VMA) mechanism. VMA computes attention scores as the volume of the parallelotope spanned by query vectors and multimodal key vectors, inherently enabling the modeling of joint dependencies among arbitrary numbers of modalities. The approach achieves significant improvements in multimodal fusion performance while maintaining computational efficiency, outperforming state-of-the-art methods across multiple benchmark tasks and demonstrating the effectiveness of explicitly modeling high-order interactions.
This work addresses the challenge in multimodal sentiment analysis where modality-specific signal refinement and cross-modal interaction modeling often interfere with each other due to conflicting optimization objectives. To resolve this, the authors propose SeRIn, a novel architecture that decouples modality separation and cross-modal interaction into structured priors through a three-stage pipeline—separation, refinement, and integration—processing unimodal representations and cross-modal interactions via independent pathways before fusing them at the prediction stage. Notably, SeRIn adaptively adjusts modality weights without requiring explicit supervision. The method achieves state-of-the-art performance on the CH-SIMS and CMU-MOSEI benchmarks, yielding significant improvements across all evaluation metrics.