KVE-KD: Key Visual Evidence-Guided Knowledge Distillation for Vision-Language Models

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing knowledge distillation methods, where uneven visual token supervision or static selection conflates task-relevant cues with background noise, thereby constraining cross-modal reasoning capabilities. To overcome this, we propose a dynamic distillation framework guided by critical visual evidence. Specifically, our approach leverages textual anchors to localize fusion layers and combines attention distributions with normalized entropy to filter task-relevant visual tokens. Furthermore, targeted feature distillation is achieved through iterative contribution removal. Extensive experiments demonstrate that the proposed method surpasses state-of-the-art approaches across six benchmarks, yielding significant improvements in complex reasoning and fine-grained understanding without introducing additional inference overhead.
📝 Abstract
Knowledge distillation is crucial for deploying vision-language models on resource-constrained devices. However, existing methods typically impose uniform supervision across visual tokens or rely on static token selection, which confuses task-relevant cues with background noise and degrades cross-modal reasoning. To address this limitation, we propose Key Visual Evidence-guided Knowledge Distillation (KVE-KD), a framework that dynamically focuses feature distillation on task-relevant visual tokens identified by the teacher model. Specifically, KVE-KD appoints the final pre-generation textual token as a unified semantic anchor and identifies the target cross-modal fusion layer by analyzing changes in the anchor representation through iterative visual-token contribution removal. Within this layer, KVE-KD ranks visual tokens via the anchor-conditioned attention distribution and selects the most informative visual tokens as key visual evidence with normalized entropy. The key visual evidence subsequently guides focused visual feature distillation, making the student align closely with the teacher's task-relevant visual representations while suppressing irrelevant background information. Extensive experiments on six benchmarks demonstrate that KVE-KD outperforms state-of-the-art cross-modal distillation methods, with particularly pronounced gains on tasks requiring complex reasoning and fine-grained visual understanding. Importantly, these improvements are achieved without introducing any inference-time overhead. The source code is available at https://github.com/zhangjianbin07/KVE-KD.
Problem

Research questions and friction points this paper is trying to address.

Knowledge Distillation
Vision-Language Models
Visual Tokens
Cross-modal Reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Knowledge Distillation
Vision-Language Models
Key Visual Evidence
Dynamic Token Selection
Cross-Modal Reasoning
🔎 Similar Papers
No similar papers found.
J
Jianbin Zhang
Faculty of Data Science, City University of Macau, SAR Macao, China
X
Xin Sun
Faculty of Data Science, City University of Macau, SAR Macao, China
S
Shanwen Wang
Faculty of Data Science, City University of Macau, SAR Macao, China
Wei Ye
Wei Ye
Tongji University
data miningmachine learningrepresentation learningdeep learning
S
Susanto Rahardja
School of Information and Electronic Engineering, Zhejiang University, Hangzhou, China