multimodal fusion

Developing methods to align and combine features across modalities and scales (e.g., attention-based fusion, vision-language alignment, heterogeneous temporal representations) so models can represent, compose, and generate grounded multimodal content.

multimodalfusion

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques

Jun 05, 2025
JA
Jisu An
🏛️ Seoul National University | University of California San Diego | Chung-Ang University

Existing research lacks a systematic analysis of modality–language backbone integration mechanisms in multimodal large language models (MLLMs). Method: We systematically survey 125 MLLMs published between 2021 and 2025, proposing the first large-language-model-centric three-dimensional taxonomy—architectural integration, representation learning, and training paradigms—to unify cross-modal alignment pathways. Leveraging bibliometric analysis, architectural decoupling, and fine-grained modeling of embeddings and loss functions, we characterize evolutionary patterns in fusion granularity, joint representation design, and objective function development. Contribution/Results: This work fills a critical theoretical gap in structured analysis of MLLM fusion mechanisms, identifies shared bottlenecks—including misaligned modality-specific feature hierarchies and suboptimal cross-modal optimization objectives—and delivers a reusable theoretical framework and practical guidelines for developing robust, scalable multimodal foundation models.

Analysis of 125 MLLMs to identify emerging patternsClassification framework for MLLMs based on key dimensionsSystematic understanding of multimodal integration with LLMs

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the lack of principled understanding in existing literature regarding the choice between cross-attention and feature concatenation strategies for multimodal fusion, which has largely relied on empirical heuristics. Through controlled experiments and theoretical analysis, we demonstrate for the first time that feature alignment quality is the key determinant of fusion strategy performance: under pre-aligned features, concatenation consistently outperforms cross-attention by 4.1–5.1 percentage points across all data scales, with its advantage becoming more pronounced as alignment degrades. Building on this insight, we develop a theoretical decision framework grounded in sample complexity and validate our findings using features extracted from ResNet-18 and CLIP ViT-B/32 on controlled datasets.

concatenationcross-attentionfeature alignment

Multimodal Representation Learning and Fusion

Jun 25, 2025
QJ
Qihang Jin
🏛️ AI Agent Lab | Vokram Group | University of Bologna | University of Minnesota | Singapore General Hospital

Multimodal learning faces critical challenges including difficulty in cross-source information fusion, poor robustness to modality missing, and vulnerability to adversarial attacks. To address these, we propose a robust multimodal representation learning framework. Methodologically, we design a contrastive learning–based cross-modal alignment mechanism with cross-attention, enabling unsupervised and self-supervised fusion; integrate AutoML-driven dynamic architecture search to enhance adaptability to incomplete inputs and adversarial perturbations; and establish a unified benchmarking framework for comprehensive evaluation. Our approach achieves significant performance gains on vision-language understanding and speech-text joint modeling tasks. Moreover, it introduces a reproducible, extensible evaluation standard system, advancing general-purpose multimodal representation paradigms. The framework demonstrates superior robustness under modality dropout and adversarial conditions while maintaining high accuracy across diverse multimodal benchmarks.

Addressing challenges like missing inputs and adversarial attacksCombining diverse data sources for better AI understandingImproving evaluation metrics for cross-domain model comparison

This study investigates whether spatial alignment in multimodal representation learning degrades modality-specific information—particularly in remote sensing fusion of heterogeneous sources (e.g., optical and SAR). We first establish a theoretical analysis framework revealing how alignment operations inherently erode modality-unique semantic content. To address this, we propose a self-supervised contrastive learning paradigm that jointly optimizes semantic alignment and modality fidelity. Extensive experiments on real-world remote sensing datasets demonstrate that aggressive spatial alignment improves cross-modal consistency but substantially compromises modality-discriminative feature representation. Our method preserves alignment performance while boosting modality-specific representation capability by 12.7% (average improvement). The work provides an interpretable trade-off principle between alignment and specificity for multimodal remote sensing fusion and releases open-source code and a benchmark dataset.

Analyzes information loss in alignment strategies for multimodal satellite dataExplores contrastive learning limitations for combining Earth observation modalitiesInvestigates whether multimodal alignment preserves modality-specific task-relevant information

Interpretation on Multi-modal Visual Fusion

Aug 19, 2023
HC
Hao Chen
🏛️ Southeast University | Beijing University of Technology

RGB-D multimodal fusion mechanisms have long suffered from poor interpretability, and the fundamental nature of cross-modal complementarity remains unclear. Method: This paper establishes the first interpretability analysis framework specifically targeting the fusion process, introducing a joint metric of semantic variance and feature similarity to systematically characterize cross-modal representation consistency, intra-modal evolutionary patterns, and collaborative optimization logic. Through cross-layer feature comparison and quantitative semantic analysis, we identify a prevalent imbalance between consistency and specificity in mainstream fusion strategies. Contribution/Results: We formalize a “specificity-driven inference under consistency constraints” principle that explicates cross-modal complementarity. Our framework provides both theoretical foundations and a verifiable evaluation paradigm for designing trustworthy, generalizable multimodal fusion models.

Analyzing feature consistency and specialty across RGB-D modalitiesDeveloping improved fusion strategies for multi-modal RGB-D learningUnderstanding complementary and fusion mechanisms in RGB-D models

Everything is a Video: Unifying Modalities through Next-Frame Prediction

Nov 15, 2024
GT
G. Thomas
🏛️ Durham University

Traditional multimodal approaches rely on modality-specific encoders and late fusion, limiting scalability and cross-modal generalization. This work proposes a unified paradigm that reformulates diverse multimodal tasks—including text, image, audio, and video processing—as “next-frame prediction,” with all inputs and outputs represented as serialized video frames, enabling end-to-end, single-model inference. It pioneers the task-redefinition strategy—previously confined to NLP—within multimodal learning, eliminating modality-specific design in favor of a fully modality-agnostic architecture. The framework integrates cross-modal tokenization, sequential frame representation, autoregressive modeling, and a shared Transformer decoder, supporting joint multi-task pretraining. Empirical evaluation demonstrates strong zero-shot and few-shot generalization across text-to-text, image-to-text, video-to-video, video-to-text, and audio-to-text tasks. Crucially, adaptation requires only lightweight task-specific heads, underscoring its flexibility and efficiency.

Enabling seamless knowledge transfer across different tasks and modalitiesOvercoming modality-specific encoder limitations in multimodal learningUnifying diverse modalities via next-frame prediction

Latest Papers

What's happening recently
View more

This paper addresses the cross-modal semantic gap in multimodal understanding through a unified alignment–translation–fusion–transfer framework. Methodologically: (1) a spatial reasoning BERT is introduced to map spatial language to 2D layouts; (2) a medical term spatial co-occurrence loss is designed to ground textual descriptions in 3D anatomical locations; (3) a structured text-to-knowledge graph fact linking benchmark with interpretability is established; and (4) a multi-stream feature fusion mechanism coupled with cross-modal knowledge distillation enables lightweight RGB-based action recognition. Key contributions include: the first spatial semantic alignment model, joint anatomical-spatial representation learning, a standardized, interpretable knowledge graph linking benchmark, and a novel unimodal distillation paradigm that achieves near-fused performance without multimodal inputs. Experiments demonstrate significant improvements across all tasks: the RGB-only model attains accuracy comparable to multimodal baselines while reducing computational overhead by over 60%.

Advances multimodal fusion for action recognition and knowledge transferEnhances machine understanding of multimodal inputs through alignment and translationImproves spatial language decoding into visual representations for scene generation

Multi-Grained Text-Guided Image Fusion for Multi-Exposure and Multi-Focus Scenarios

Dec 23, 2025
MT
Mingwei Tang
🏛️ Xidian University | Nanyang Technological University | Xi'an University of Technology

To address the challenge of jointly modeling dynamic range and depth-of-field variations in multi-exposure and multi-focus image fusion, this paper proposes the first hierarchical text-guided fusion framework. Methodologically, we design a multi-granularity text encoder—capturing fine-grained details, mid-granularity structures, and coarse-granularity semantics—and build a hierarchical cross-modulation network. We further introduce a granularity-aware supervised loss and a saliency-driven semantic enhancement module to achieve precise cross-modal feature alignment and adaptive modulation. Our key innovation lies in explicitly embedding text granularity priors into the fusion process during training, eliminating the need for textual input at test time. Extensive experiments on mainstream multi-exposure and multi-focus benchmarks demonstrate consistent superiority over state-of-the-art methods, with significant improvements in PSNR and SSIM. The framework exhibits strong generalization across diverse fusion scenarios.

Addresses disparities in dynamic range and focus depth between imagesEnhances fusion with multi-grained text guidance and cross-modal alignmentSynthesizes high-quality images from varied exposure and focus inputs

This work addresses the challenge in multimodal sentiment analysis where modality-specific signal refinement and cross-modal interaction modeling often interfere with each other due to conflicting optimization objectives. To resolve this, the authors propose SeRIn, a novel architecture that decouples modality separation and cross-modal interaction into structured priors through a three-stage pipeline—separation, refinement, and integration—processing unimodal representations and cross-modal interactions via independent pathways before fusing them at the prediction stage. Notably, SeRIn adaptively adjusts modality weights without requiring explicit supervision. The method achieves state-of-the-art performance on the CH-SIMS and CMU-MOSEI benchmarks, yielding significant improvements across all evaluation metrics.

cross-modal interactionmodality-specific refinementmultimodal fusion

This work proposes AlignMamba-2, a novel framework addressing key challenges in multimodal sentiment analysis—namely, cross-modal alignment difficulty, modality heterogeneity, and computational inefficiency. AlignMamba-2 integrates a dual-alignment mechanism based on optimal transport and maximum mean discrepancy to enhance cross-modal consistency. It further introduces a modality-aware Mixture-of-Experts architecture that effectively combines modality-specific and shared experts to better model heterogeneous data. Leveraging the efficient Mamba backbone, the proposed method achieves state-of-the-art performance on four benchmark datasets—CMU-MOSI, CMU-MOSEI, NYU-Depth V2, and MVSA-Single—demonstrating significant improvements over existing approaches in both accuracy and inference efficiency.

computational efficiencycross-modal alignmentmodality heterogeneity

This work addresses the suboptimal performance in multimodal representation alignment caused by modality gaps and data scarcity. To this end, the authors propose a disentangled representation learning framework based on shared and modality-specific codebooks. Leveraging a compositional vector quantization mechanism, the method decomposes multimodal features into shared semantic components and modality-unique components, and employs a progressive alignment strategy to optimize the alignment space without requiring fully paired data. The unified shared codebook effectively bridges the modality gap, while the modality-specific codebooks mitigate dominant-modality bias, enabling more balanced multimodal fusion. The approach achieves state-of-the-art performance across classification and retrieval tasks spanning nine modalities, including text, images, video, and audio.

cross-modal discrepancydata scarcitymodality-unique features

Hot Scholars

GZ

Guangtao Zhai

Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI Evaluation
ZL

Ziwei Liu

Associate Professor, Nanyang Technological University
Computer VisionMachine LearningComputer Graphics
PW

Pengfei Wan

Head of Kling Video Generation Models, Kuaishou Technology
Generative ModelsComputer VisionMultimodal AIComputer Graphics
MH

Ming-Hsuan Yang

University of California at Merced; Google DeepMind
Computer VisionMachine LearningArtificial Intelligence