Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models

๐Ÿ“… 2024-12-26
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Large vision-language models (LVLMs) overly rely on deepest-layer visual features while neglecting complementary information across intermediate layers. Method: We propose an instruction-guided dynamic visual feature aggregation mechanismโ€”the first of its kind to enable instruction-driven, adaptive selection, weighting, and cross-layer interaction of multi-depth visual features without increasing the number of visual tokens. Leveraging task-aware attention aggregation and instruction-conditioned gating, the method jointly optimizes fine-grained perception (via low-level features) and semantic understanding (via mid- to high-level features). Contribution/Results: Our approach achieves significant performance gains across 18 benchmarks spanning six diverse vision-language tasks, demonstrating both the effectiveness and strong generalizability of dynamic, hierarchical visual feature utilization.

Technology Category

Computer Vision: Large Vision ModelsMachine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Language Grounding & Multi-modal NLP

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
๐Ÿ“ Abstract
Large Vision-Language Models (LVLMs) have achieved remarkable success in a wide range of multimodal tasks by integrating pre-trained vision encoders and large language models. However, current LVLMs primarily rely on visual features extracted from the final layers of the vision encoder, overlooking the complementary information available in shallower layers. While recent approaches have explored the use of multilayer visual features in LVLMs, they tend to be task-agnostic and fail to examine the dependencies of hierarchical visual features on specific tasks. To address these gaps, we systematically investigate the contributions of visual features from different encoder layers using 18 benchmarks spanning 6 task categories. Our findings reveal that multilayer features provide complementary strengths with varying task dependencies, and uniform fusion leads to suboptimal performance. Building on these insights, we propose the instruction-guided vision aggregator, a module that dynamically integrates multi-layer visual features based on textual instructions, without increasing the number of visual tokens. Extensive evaluations demonstrate the superior performance of our method. Additionally, an in-depth analysis of the aggregator's behavior highlights the dominance of mid-to-high-level features in semantic-rich tasks and the critical role of low-level features in fine-grained perception.
Problem

Research questions and friction points this paper is trying to address.

Visual Language Models
Multilevel Image Information
Task-specific Adaptation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Task-aware Multilevel Information Fusion
Visual Language Models Optimization
Selective Utilization of Image Layers
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
X
Xu Li
Institute of Big Data, Fudan University, 220 Handan Road, Shanghai, 200433, China
Y
Yi Zheng
School of Computer Science, Fudan University, 220 Handan Road, Shanghai, 200433, China
Haotian Chen
Haotian Chen
University of California, Los Angeles
Political EconomyNon-market StrategyAmerican Politics
X
Xiaolei Chen
Institute of Big Data, Fudan University, 220 Handan Road, Shanghai, 200433, China
Yuxuan Liang
Yuxuan Liang
Assistant Professor, Hong Kong University of Science and Technology (Guangzhou)
Spatio-Temporal Data MiningUrban ComputingUrban AIFoundation ModelsTime Series
C
Chenghang Lai
Institute of Big Data, Fudan University, 220 Handan Road, Shanghai, 200433, China
B
Bin Li
School of Computer Science, Fudan University, 220 Handan Road, Shanghai, 200433, China
Xiangyang Xue
Xiangyang Xue
Professor of Computer Science, Fudan University
Computer VisionPattern RecognitionMachine Learning