Distill the Visual Evidence, Not Just the Answer: Cross-World On-Policy Distillation for Vision-Language Models

πŸ“… 2026-09-29
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitation of existing vision-language model (VLM) distillation, which supervises only outputs and causes student models to fit answers while disregarding genuine visual evidence. To this end, we propose Cross-World Online Policy Distillation, which constructs differentiated visual scenarios to explicitly optimize the student’s responsiveness to changes in visual evidence. A cross-world belief transfer mechanism is further introduced to enforce reasoning grounded in authentic visual cues. Additionally, we establish CWBench, a benchmark for systematically diagnosing visual consistency. Experiments demonstrate that our method yields an average performance improvement of 1.2 points. Notably, a 4B-parameter student model surpasses its 552B-parameter teacher by 22.4 points on CWBench, effectively achieving reliable transfer of visual understanding capabilities.
πŸ“ Abstract
A central goal of vision-language model (VLM) distillation is to transfer both the teacher's language capabilities and its visual understanding. However, existing methods primarily supervise the student's output, leaving visual understanding implicit. Our analysis reveals that a student can match the teacher's answer without relying on the same visual evidence, raising the question: how can we ensure the student responds to the visual information that actually determines the answer? To this end, we propose \textbf{Cross-World On-Policy Distillation (CW-OPD)}, which explicitly supervises the student's response to changes in visual evidence. For each example, CW-OPD constructs two visual worlds that share the question and scene context but differ in answer-critical evidence, yielding different answers. We perform on-policy distillation in both worlds and distill the teacher's cross-world belief transition, encouraging the student to match not only \emph{what} the teacher predicts but also \emph{why} its prediction changes with the evidence. A gradient analysis shows that this term is invariant to errors shared by both worlds and supplies a corrective signal invisible to endpoint matching alone. In this way, CW-OPD makes reliance on the relevant visual evidence an explicit distillation target rather than an implicit consequence of output matching. To diagnose whether a model truly grounds its answers in visual evidence, we introduce CWBench, which measures cross-world consistency via Cross-World Pair Accuracy (CWPA). Experiments on Qwen3.5-4B show that CW-OPD outperforms the strongest baseline by \textbf{1.2} points on average, and the 4B student exceeds DeepSeek-V4.1 (552B) by \textbf{22.4} CWPA points on CWBench. Code is released in https://github.com/baokou-fw2/CWAD.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Knowledge Distillation
Visual Evidence
Visual Grounding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-World On-Policy Distillation
Vision-Language Models
Visual Evidence
Knowledge Distillation
CWBench
πŸ”Ž Similar Papers
No similar papers found.
Y
Yuanhao Sun
Shanghai Jiao Tong University, Shanghai, China
H
Huawei Ji
Shanghai Jiao Tong University, Shanghai, China
Jiaxin Ding
Jiaxin Ding
Shanghai Jiao Tong University
Spatio-temporal Data MiningReinforcement LearningLarge Language Model Reasoning
L
Luoyi Fu
Shanghai Jiao Tong University, Shanghai, China
X
Xinbing Wang
Shanghai Jiao Tong University, Shanghai, China