deep feature fusion

Design and implement trainable architectures and modules that combine heterogeneous feature representations—e.g., multiple modalities (images, text, metadata), parallel model streams, or task-specific embeddings—into unified joint features or decision layers. This includes building learnable fusion blocks (addition-based, shared/joint, or mixed), choosing data- vs. model-level fusion, integrating fusion into classification or classification–matching pipelines, and linking multimodal encoders with larger models (e.g., LLMs) or agent-mediated workflows for downstream inference.

deepfeaturefusion

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.57
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Multimodal Representation Learning and Fusion

Jun 25, 2025
QJ
Qihang Jin
🏛️ AI Agent Lab | Vokram Group | University of Bologna | University of Minnesota | Singapore General Hospital

Multimodal learning faces critical challenges including difficulty in cross-source information fusion, poor robustness to modality missing, and vulnerability to adversarial attacks. To address these, we propose a robust multimodal representation learning framework. Methodologically, we design a contrastive learning–based cross-modal alignment mechanism with cross-attention, enabling unsupervised and self-supervised fusion; integrate AutoML-driven dynamic architecture search to enhance adaptability to incomplete inputs and adversarial perturbations; and establish a unified benchmarking framework for comprehensive evaluation. Our approach achieves significant performance gains on vision-language understanding and speech-text joint modeling tasks. Moreover, it introduces a reproducible, extensible evaluation standard system, advancing general-purpose multimodal representation paradigms. The framework demonstrates superior robustness under modality dropout and adversarial conditions while maintaining high accuracy across diverse multimodal benchmarks.

Addressing challenges like missing inputs and adversarial attacksCombining diverse data sources for better AI understandingImproving evaluation metrics for cross-domain model comparison

Current biomedical multimodal models face critical bottlenecks: reliance on end-to-end training, exponential growth in computational complexity with modality count, severe performance degradation under extreme modality imbalance, and rigid topological coupling. To address these, we propose MM-Lego—a tuning-free, universal multimodal fusion framework. It introduces a novel frequency-domain feature harmonization mechanism that achieves shape alignment and interference-free merging of arbitrary unimodal encoders. We further design modality-agnostic wrappers and zero-/few-shot model merging strategies, enabling topology-agnostic fusion and robust modeling under highly imbalanced modalities. Crucially, MM-Lego requires no fine-tuning yet matches or surpasses end-to-end models in performance, while maintaining full encoder compatibility. Evaluated across seven biomedical benchmark datasets, it achieves state-of-the-art results on five—demonstrating unprecedented flexibility, efficiency, and generalizability in biomedical multimodal learning.

Enabling flexible model merging with minimal fine-tuningHandling diverse data modalities in biomedical machine learningOvercoming limitations of existing multimodal fusion approaches

Multimodal fusion faces challenges in adapting heterogeneous data (e.g., images, text, audio) and incurs high costs in manually designing effective architectures. Method: This paper proposes SAMAS, a sampling-driven multimodal mixer architecture search framework. SAMAS introduces the first end-to-end joint optimization paradigm for multimodal learning, simultaneously searching for optimal MLP-based mixer structures, modality-specific encoder combinations, and fusion functions. It employs a lightweight micro-benchmark–based sampling evaluation mechanism to accelerate architecture assessment. Contribution/Results: SAMAS achieves 3–5× higher search efficiency than reinforcement learning or evolutionary algorithms. It attains state-of-the-art fusion performance across multiple standard multimodal benchmarks—including MM-IMDB, CMU-MOSEI, and UR-FUNNY—while substantially reducing reliance on manual architectural design and lowering computational overhead.

Deep Learning MethodologyMulti-modal LearningOptimal Network Structure

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques

Jun 05, 2025
JA
Jisu An
🏛️ Seoul National University | University of California San Diego | Chung-Ang University

Existing research lacks a systematic analysis of modality–language backbone integration mechanisms in multimodal large language models (MLLMs). Method: We systematically survey 125 MLLMs published between 2021 and 2025, proposing the first large-language-model-centric three-dimensional taxonomy—architectural integration, representation learning, and training paradigms—to unify cross-modal alignment pathways. Leveraging bibliometric analysis, architectural decoupling, and fine-grained modeling of embeddings and loss functions, we characterize evolutionary patterns in fusion granularity, joint representation design, and objective function development. Contribution/Results: This work fills a critical theoretical gap in structured analysis of MLLM fusion mechanisms, identifies shared bottlenecks—including misaligned modality-specific feature hierarchies and suboptimal cross-modal optimization objectives—and delivers a reusable theoretical framework and practical guidelines for developing robust, scalable multimodal foundation models.

Analysis of 125 MLLMs to identify emerging patternsClassification framework for MLLMs based on key dimensionsSystematic understanding of multimodal integration with LLMs

Multi-Layer Visual Feature Fusion in Multimodal LLMs: Methods, Analysis, and Best Practices

Mar 08, 2025
JL
Junyan Lin
🏛️ Ocean University of China | Zhejiang Gongshang University | Genmo.ai | Meituan Inc. | NUS

Systematic investigation remains lacking on effective multimodal fusion of hierarchical visual features in multimodal large language models (MLLMs), particularly regarding optimal visual layer selection and fusion paradigms with the language model. Method: We extract multilevel visual features from CLIP/ViT, and systematically evaluate fusion strategies—including learnable weighting, concatenation, and attention-based fusion—alongside ablation-driven layer importance assessment for modular integration. Contribution/Results: Our empirical study is the first to reveal that cross-stage (e.g., early + late) visual feature fusion significantly improves generalization, whereas intra-stage stacking degrades performance; input-side direct concatenation emerges as the most stable and efficient fusion paradigm. On benchmarks including MMBench and OCRBench, our approach achieves an average accuracy gain of 2.3% and improves fusion stability by 37%. The code is open-sourced and has become a new de facto standard for visual feature fusion in MLLMs.

Analyzes impact of multi-layer visual feature integration on performance.Explores optimal layer selection for visual features in MLLMs.Identifies best fusion strategies between visual and language models.

Latest Papers

What's happening recently
View more

This work addresses the limitations of existing approaches that predominantly rely on non-native late-fusion architectures and lack a systematic definition or unified framework for native multimodal modeling. The paper proposes a formal roadmap for Native Multimodal Modeling (NMM), introducing, for the first time, a rigorous formulation of “architectural nativeness.” Leveraging an input–output duality perspective, it categorizes models into three types: Multi-to-Text, Multi-to-Target, and Multi-to-Multi. A full-stack, industrial-grade NMM framework is developed, encompassing data governance, early/mid-stage fusion, end-to-end training, and deployment, all realized through a unified Transformer architecture that enables symbiotic cross-modal understanding and generation. This paradigm demonstrates superior performance and strong scalability across multiple tasks, offering a clear pathway toward truly native multimodal models.

architectural nativitymultimodal architecturemultimodal fusion

Existing model-serving frameworks struggle to efficiently support increasingly complex composite multimodal models. To address this challenge, this work proposes M*, a novel system that introduces a modular Walk Graph abstraction to uniformly represent composite AI models as dataflow graphs. This abstraction enables flexible composition of arbitrary model components, cluster deployment, and model-agnostic distributed runtime optimizations. By leveraging a graph traversal mechanism, M* efficiently handles cross-modal, multitask requests, achieving significant performance gains: it reduces end-to-end latency by 20% over vLLM-Omni in text-to-image generation, improves real-time factor by 2.9× and throughput by 2.7× in text-to-speech tasks, and accelerates robot planning workloads by up to 12.5×.

architectural diversitycomposite architecturesmodel serving

This work systematically investigates the core mechanisms of modality interaction in multimodal pretraining, focusing on knowledge transfer, synergistic effects, and fusion timing. Through experiments on both synthetic and large-scale real-world datasets, it provides the first empirical evidence of asymmetric cross-modal knowledge flow and demonstrates that data complexity governs whether modalities exhibit synergy or competition. The study further validates that early unified fusion consistently outperforms late alignment. Leveraging an architecture featuring shared attention and normalization layers with modality-specific feedforward components, the proposed approach is evaluated on a 13.5B mixture-of-experts model trained on 2 trillion tokens, confirming its effectiveness. Additionally, the paper introduces a highly efficient pretraining strategy that achieves strong generative performance using only 5% of the typical computational budget.

early unificationknowledge flowmodality interaction

Existing unified multimodal models (UMMs) lack a fair evaluation framework due to disparities in architecture, training paradigms, and implementation details. This work proposes the first unified codebase built on PyTorch that supports multimodal understanding, generation, and editing tasks, offering compatibility with diverse backbone architectures, model scales, and datasets. The project introduces standardized evaluation protocols and a consistent interface, enabling—for the first time—fair and reproducible comparisons across heterogeneous UMMs. Furthermore, it integrates a multidimensional benchmarking suite assessing perceptual quality, reasoning, compositionality, and instruction-following capabilities, along with a streamlined post-training pipeline. This comprehensive infrastructure lays a foundational framework for developing more powerful and generalizable unified multimodal systems.

evaluation frameworkmodel heterogeneitymultimodal understanding

Existing approaches struggle to effectively model the dynamic interplay among redundant, unique, and synergistic information at the sample level in multimodal learning. This work proposes a novel information-theoretic decomposition–based paradigm for multimodal interaction learning, offering the first systematic analysis of the importance of sample-level interactions. By employing a variational architecture, the method explicitly disentangles these three interaction components and integrates a component-aware fine-tuning strategy to adaptively leverage them. Extensive experiments across diverse tasks and architectures consistently demonstrate the superiority of the proposed approach over current state-of-the-art methods, confirming its effectiveness and generality in finely modeling sample-level multimodal interactions.

dynamic interactioninformation decompositionmultimodal interaction

Hot Scholars

JT

Jin Tang

Anhui University
Computer visionintelligent video analysis
YW

Yaowei Wang

The Hong Kong Polytechnic University
JH

Jianxi Huang

Professor in China Agricultural University
Data assimilationClimate changeAgricultural remote sensingCrop modeling with remote sensing data assimilation
MZ

Moyu Zhang

Beijing University of Posts and Telecommunications、Alibaba Group
Knowledge TracingInformation RetrievalRecommender System