Score
Designs, implements, and evaluates models and systems that jointly represent, align, and reason over multiple data modalities (e.g., text, images, audio, video), including modality-specific encoders, cross‑modal fusion and attention mechanisms, and grounding/alignment methods. Analyzes these models’ performance, robustness to missing or noisy modalities, cross‑modal retrieval and reasoning capabilities, and the interpretability of their multimodal representations.
This paper systematically analyzes four core challenges in multimodal alignment and fusion—cross-modal misalignment, semantic modality gap, computational bottlenecks, and data noise/heterogeneity—that hinder model robustness, generalizability, and scalability. To address them, the authors propose a unified taxonomy covering 200+ studies, revealing an integrated “alignment–fusion” co-design paradigm. They further introduce three principled solution pathways: (1) noise-robust learning, (2) heterogeneous representation modeling, and (3) few-shot cross-modal transfer—unifying contrastive learning, cross-modal attention, latent-space alignment, graph neural networks, and differentiable architecture search. Extensive experiments on social media analysis, medical imaging, and sentiment recognition validate the efficacy of the framework. The work distills actionable design principles for scalable, robust, and generalizable multimodal learning, establishing a new benchmark for both theoretical research and industrial deployment.
This paper addresses the challenges in developing general-purpose multimodal large language models (MLLMs) capable of cross-modal coordination across six generative modalities: text, image, music, video, human motion, and 3D objects. To this end, it proposes a novel unified architecture integrating Transformer-based and diffusion-based paradigms, augmented with self-supervised learning (SSL), mixture-of-experts (MoE), reinforcement learning from human feedback (RLHF), and chain-of-thought (CoT) reasoning. The work introduces the first taxonomy covering all six modalities and identifies shared enabling mechanisms for cross-modal transfer. It further argues that structured reasoning and modular decoupling are critical to improving interpretability and generalization. The resulting comprehensive MLLM technology landscape clarifies common bottlenecks and transferable methodologies, providing both theoretical foundations and practical guidelines for building universal, adaptive, and interpretable multimodal systems.
Cross-modal retrieval suffers from semantic misalignment due to geometric inconsistency between image and text representations. This work formulates image–text matching as a metric alignment problem in the embedding space. We systematically demonstrate that: (1) the Wasserstein distance quantitatively characterizes inter-modal distributional discrepancy; (2) cosine similarity exhibits superior robustness over Euclidean distance, KL divergence, and other conventional metrics for alignment; and (3) standard MLPs inadequately capture complex cross-modal interactions. To address these issues, we propose a lightweight neural metric learning framework that integrates vision-language models with unimodal encoders, jointly optimizing four classical metrics and two learnable neural metrics. Our approach achieves significant improvements in Recall@K across standard cross-modal retrieval benchmarks. Moreover, it provides practical, deployment-oriented alignment evaluation criteria and architectural design guidelines.
This study investigates cross-modal decision-making mechanisms in vision-language models (VLMs) under modality conflicts—e.g., an image of a dog paired with the caption “This is a cat.” We systematically construct conflict-rich multimodal samples and employ attention head localization, representation space analysis, and instruction-guided modality selection tasks. Our analysis reveals an intrinsic modality bias in VLMs and identifies two functionally distinct architectural components: (i) dedicated attention heads that regulate modality preference, and (ii) transferable, modality-agnostic “router heads” that dynamically route information across modalities. Crucially, targeted intervention on router heads significantly improves model accuracy in detecting multimodal consistency. This work provides the first empirical evidence of a hierarchical modality fusion architecture within VLMs, uncovering interpretable, controllable mechanisms for multimodal reasoning and cross-modal alignment.
RGB-D multimodal fusion mechanisms have long suffered from poor interpretability, and the fundamental nature of cross-modal complementarity remains unclear. Method: This paper establishes the first interpretability analysis framework specifically targeting the fusion process, introducing a joint metric of semantic variance and feature similarity to systematically characterize cross-modal representation consistency, intra-modal evolutionary patterns, and collaborative optimization logic. Through cross-layer feature comparison and quantitative semantic analysis, we identify a prevalent imbalance between consistency and specificity in mainstream fusion strategies. Contribution/Results: We formalize a “specificity-driven inference under consistency constraints” principle that explicates cross-modal complementarity. Our framework provides both theoretical foundations and a verifiable evaluation paradigm for designing trustworthy, generalizable multimodal fusion models.
This study addresses the robustness deficiency of multimodal large language models (MLLMs) under modality conflicts—such as audio-visual inconsistencies or text-based misdirection—revealing their over-reliance on single modalities. To systematically evaluate this vulnerability, we introduce MMA-Bench, the first benchmark explicitly designed for modality conflict assessment. Our method proposes a context-aware modality alignment fine-tuning framework: it constructs test sets using video–task pairs, integrates black-box and white-box interpretability analyses, and designs a modality alignment loss to dynamically calibrate cross-modal attention weights during inference. Empirical results demonstrate substantial improvements in reasoning accuracy and multimodal grounding under contradictory inputs across diverse open- and closed-source MLLMs. The approach generalizes effectively across architectures and establishes a new paradigm for robust multimodal reasoning.
Traditional multimodal approaches rely on modality-specific encoders and late fusion, limiting scalability and cross-modal generalization. This work proposes a unified paradigm that reformulates diverse multimodal tasks—including text, image, audio, and video processing—as “next-frame prediction,” with all inputs and outputs represented as serialized video frames, enabling end-to-end, single-model inference. It pioneers the task-redefinition strategy—previously confined to NLP—within multimodal learning, eliminating modality-specific design in favor of a fully modality-agnostic architecture. The framework integrates cross-modal tokenization, sequential frame representation, autoregressive modeling, and a shared Transformer decoder, supporting joint multi-task pretraining. Empirical evaluation demonstrates strong zero-shot and few-shot generalization across text-to-text, image-to-text, video-to-video, video-to-text, and audio-to-text tasks. Crucially, adaptation requires only lightweight task-specific heads, underscoring its flexibility and efficiency.
Conventional multimodal learning assumes explicit representation alignment is universally beneficial, yet its causal impact remains unverified. Method: This work systematically investigates the conditional effects of enforced alignment, proposing a controllable contrastive learning module that dynamically modulates alignment strength and establishes a quantitative relationship between alignment strength and inter-modal information redundancy—derived via information decomposition and synthetic data modeling. Contribution/Results: We demonstrate that alignment efficacy is not universal: strong alignment improves performance under high redundancy but harms modality-specific representation learning under low redundancy. To address this, we introduce a balanced mechanism that jointly preserves shared semantics and modality-specific characteristics. Empirical evaluation on both synthetic and real-world benchmarks confirms that this mechanism significantly enhances the generalization capability of unimodal encoders. Our core contribution is the formal establishment of “redundancy-dependent alignment”—a principled, interpretable, and tunable paradigm for multimodal representation learning.
This study investigates whether vision-language models genuinely rely on visual information for reasoning. To this end, we introduce CrossMath, the first multimodal mathematical reasoning benchmark enabling rigorous cross-modal comparison, with human-verified alignment among text-only, image-only, and combined text-image questions. Systematic evaluation reveals that current models achieve their best performance on text-only inputs, while the inclusion of images often degrades accuracy—highlighting a significant deficiency in their visual reasoning capabilities. Fine-tuning on CrossMath substantially enhances model performance across all modalities and generalizes to broader visual reasoning tasks.
Current multimodal large language model (MLLM) evaluation benchmarks suffer from a prevalence of “shortcut questions” that can be answered using only a single modality, undermining the reliable and efficient assessment of genuine cross-modal reasoning capabilities. To address this, this work proposes the Multimodal Multidimensional Item Response Theory framework (M3IRT), which extends classical Item Response Theory (IRT) to the multimodal setting by decoupling model ability and item difficulty into three distinct dimensions: visual, textual, and cross-modal. M3IRT enables precise modeling of cross-modal reasoning and effectively identifies and filters out shortcut questions. Experiments across three benchmarks with 24 models demonstrate that M3IRT can extract compact, high-quality evaluation subsets from datasets containing up to 50% low-quality items, significantly improving assessment efficiency and reliability while preserving rank consistency among models.
This study addresses the challenge that multimodal models struggle to reliably bind textual prompts (e.g., “image”) to their corresponding input modalities, resulting in inadequate source modality tracking. The work formally defines and empirically investigates the “source modality monitoring” problem for the first time, introducing an evaluation paradigm grounded in target-modality information retrieval. By integrating syntactic manipulation with semantic perturbation, the authors systematically assess binding mechanisms across eleven prominent vision-language models. Their findings reveal that semantic cues dominate the binding process when modality distributions exhibit significant divergence, consistently outweighing syntactic signals. These insights offer critical evidence and a novel perspective for enhancing the reliability and robustness of multimodal agents.
This paper addresses the cross-modal semantic gap in multimodal understanding through a unified alignment–translation–fusion–transfer framework. Methodologically: (1) a spatial reasoning BERT is introduced to map spatial language to 2D layouts; (2) a medical term spatial co-occurrence loss is designed to ground textual descriptions in 3D anatomical locations; (3) a structured text-to-knowledge graph fact linking benchmark with interpretability is established; and (4) a multi-stream feature fusion mechanism coupled with cross-modal knowledge distillation enables lightweight RGB-based action recognition. Key contributions include: the first spatial semantic alignment model, joint anatomical-spatial representation learning, a standardized, interpretable knowledge graph linking benchmark, and a novel unimodal distillation paradigm that achieves near-fused performance without multimodal inputs. Experiments demonstrate significant improvements across all tasks: the RGB-only model attains accuracy comparable to multimodal baselines while reducing computational overhead by over 60%.