Remodeling Semantic Relationships in Vision-Language Fine-Tuning

πŸ“… 2025-11-11
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Existing vision-language fine-tuning methods often neglect image-internal semantic relationships emphasized by text, leading to suboptimal cross-modal alignment. To address this, we propose a semantics- and relation-driven multimodal alignment framework. First, a multi-level visual encoder explicitly models fine-grained intra-image semantic relations. Second, we design a transferable cross-attention mechanism that dynamically filters low-relevance vision–text feature pairs at the global level, enabling robust multimodal fusion. Third, semantic grouping projection is introduced to enhance cross-modal interaction. Our framework demonstrates strong generalizability, validated across eight mainstream foundation models. It achieves significant improvements over state-of-the-art methods on both visual question answering and image captioning benchmarks. Empirical results underscore the critical role of explicit intra-image relational modeling in enhancing the quality of cross-modal alignment.

Technology Category

Computer Vision: Multi-modal VisionMachine Learning: Multimodal LearningNatural Language Processing: Language Grounding & Multi-modal NLP

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
πŸ“ Abstract
Vision-language fine-tuning has emerged as an efficient paradigm for constructing multimodal foundation models. While textual context often highlights semantic relationships within an image, existing fine-tuning methods typically overlook this information when aligning vision and language, thus leading to suboptimal performance. Toward solving this problem, we propose a method that can improve multimodal alignment and fusion based on both semantics and relationships.Specifically, we first extract multilevel semantic features from different vision encoder to capture more visual cues of the relationships. Then, we learn to project the vision features to group related semantics, among which are more likely to have relationships. Finally, we fuse the visual features with the textual by using inheritable cross-attention, where we globally remove the redundant visual relationships by discarding visual-language feature pairs with low correlation. We evaluate our proposed method on eight foundation models and two downstream tasks, visual question answering and image captioning, and show that it outperforms all existing methods.
Problem

Research questions and friction points this paper is trying to address.

Improving multimodal alignment through semantic relationship modeling
Addressing overlooked semantic relationships in vision-language fine-tuning
Enhancing visual question answering and image captioning performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Extracts multilevel semantic features from vision encoders
Projects vision features to group related semantic relationships
Fuses features using inheritable cross-attention removing redundancy
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
X
Xiangyang Wu
Hangzhou International Innovation Institute, Beihang University
L
Liu Liu
School of Artificial Intelligence, Beihang University
Baosheng Yu
Baosheng Yu
Assistant Professor, Nanyang Technological University
Machine LearningDeep LearningComputer VisionAI for Medicine
J
Jiayan Qiu
University of Leicester
Z
Zhenwei Shi
Nanyang Technological University