🤖 AI Summary
Systematic investigation remains lacking on effective multimodal fusion of hierarchical visual features in multimodal large language models (MLLMs), particularly regarding optimal visual layer selection and fusion paradigms with the language model. Method: We extract multilevel visual features from CLIP/ViT, and systematically evaluate fusion strategies—including learnable weighting, concatenation, and attention-based fusion—alongside ablation-driven layer importance assessment for modular integration. Contribution/Results: Our empirical study is the first to reveal that cross-stage (e.g., early + late) visual feature fusion significantly improves generalization, whereas intra-stage stacking degrades performance; input-side direct concatenation emerges as the most stable and efficient fusion paradigm. On benchmarks including MMBench and OCRBench, our approach achieves an average accuracy gain of 2.3% and improves fusion stability by 37%. The code is open-sourced and has become a new de facto standard for visual feature fusion in MLLMs.
📝 Abstract
Multimodal Large Language Models (MLLMs) have made significant advancements in recent years, with visual features playing an increasingly critical role in enhancing model performance. However, the integration of multi-layer visual features in MLLMs remains underexplored, particularly with regard to optimal layer selection and fusion strategies. Existing methods often rely on arbitrary design choices, leading to suboptimal outcomes. In this paper, we systematically investigate two core aspects of multi-layer visual feature fusion: (1) selecting the most effective visual layers and (2) identifying the best fusion approach with the language model. Our experiments reveal that while combining visual features from multiple stages improves generalization, incorporating additional features from the same stage typically leads to diminished performance. Furthermore, we find that direct fusion of multi-layer visual features at the input stage consistently yields superior and more stable performance across various configurations. We make all our code publicly available: https://github.com/EIT-NLP/Layer_Select_Fuse_for_MLLM.