Multi-Layer Visual Feature Fusion in Multimodal LLMs: Methods, Analysis, and Best Practices

📅 2025-03-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Systematic investigation remains lacking on effective multimodal fusion of hierarchical visual features in multimodal large language models (MLLMs), particularly regarding optimal visual layer selection and fusion paradigms with the language model. Method: We extract multilevel visual features from CLIP/ViT, and systematically evaluate fusion strategies—including learnable weighting, concatenation, and attention-based fusion—alongside ablation-driven layer importance assessment for modular integration. Contribution/Results: Our empirical study is the first to reveal that cross-stage (e.g., early + late) visual feature fusion significantly improves generalization, whereas intra-stage stacking degrades performance; input-side direct concatenation emerges as the most stable and efficient fusion paradigm. On benchmarks including MMBench and OCRBench, our approach achieves an average accuracy gain of 2.3% and improves fusion stability by 37%. The code is open-sourced and has become a new de facto standard for visual feature fusion in MLLMs.

Technology Category

Computer Vision: Multi-modal VisionMachine Learning: Large Multimodal Models (LMMs)Intelligent Robots: Multimodal Perception & Sensor Fusion

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Multimodal Large Language Models (MLLMs) have made significant advancements in recent years, with visual features playing an increasingly critical role in enhancing model performance. However, the integration of multi-layer visual features in MLLMs remains underexplored, particularly with regard to optimal layer selection and fusion strategies. Existing methods often rely on arbitrary design choices, leading to suboptimal outcomes. In this paper, we systematically investigate two core aspects of multi-layer visual feature fusion: (1) selecting the most effective visual layers and (2) identifying the best fusion approach with the language model. Our experiments reveal that while combining visual features from multiple stages improves generalization, incorporating additional features from the same stage typically leads to diminished performance. Furthermore, we find that direct fusion of multi-layer visual features at the input stage consistently yields superior and more stable performance across various configurations. We make all our code publicly available: https://github.com/EIT-NLP/Layer_Select_Fuse_for_MLLM.
Problem

Research questions and friction points this paper is trying to address.

Explores optimal layer selection for visual features in MLLMs.
Identifies best fusion strategies between visual and language models.
Analyzes impact of multi-layer visual feature integration on performance.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Systematic investigation of multi-layer visual feature fusion
Optimal selection and fusion strategies for visual layers
Direct fusion at input stage enhances model performance
💼 Related Jobs
No related jobs found.
J
Junyan Lin
Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative, Institute of Digital Twin, EIT; Ocean University of China
H
Haoran Chen
Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative, Institute of Digital Twin, EIT; Zhejiang Gongshang University
Y
Yue Fan
Genmo.ai
Y
Yingqi Fan
Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative, Institute of Digital Twin, EIT
X
Xin Jin
Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative, Institute of Digital Twin, EIT
H
Hui Su
Meituan Inc.
Jinlan Fu
Jinlan Fu
National University of Singapore
Natural Language ProcessingVision and LanguageLarge Language Model
Xiaoyu Shen
Xiaoyu Shen
Eastern Institute of Technology, Ningbo
language modelmulti-modal learningreasoning