🤖 AI Summary
In multimodal policy distillation, the corrective signals provided by the teacher often conflate visual evidence, linguistic priors, and model biases, making it difficult to isolate corrections genuinely grounded in visual input. This work proposes Visual Attribution Distillation (VAD), which introduces a counterfactual intervention mechanism: by comparing the teacher’s outputs with and without visual input at the student’s prefix, VAD estimates the component of the correction attributable solely to vision and uses this to reconstruct the supervision target. Integrating centered log-probability differences, projection-based decomposition, and weakly regularized privileged-teacher supervision, VAD significantly outperforms both direct privileged-view distillation and vision-weighting baselines across six fine-grained visual benchmarks, demonstrating that the extracted visual attribution component carries rich task-relevant signals capable of effectively guiding student learning.
📝 Abstract
Multimodal on-policy distillation (OPD) transfers fine-grained visual knowledge by supervising student-generated trajectories with a privileged-view teacher. Yet its next-token corrections are source-mixed, combining visual signals with linguistic priors and teacher-specific effects. The key challenge is to estimate which corrections are supported by visual evidence, not merely where or how strongly to distill. We introduce Visual Attribution Distillation (VAD), a counterfactual target-reconstruction algorithm that estimates the visually attributable part of a teacher correction. At each student-generated prefix, VAD evaluates the same fixed teacher with the relevant evidence present and removed. The corresponding change in centered log-probabilities defines ut, a signed proxy for the visual evidence direction that estimates how revealing the evidence supports or refutes candidate tokens. VAD projects the original correction onto this proxy to obtain an intervention-aligned component and a proxy-unexplained residual, then reconstructs a student-anchored target from the former. During training, this reconstructed target supplies the primary supervision signal, while the privileged teacher contributes a weak regularizer. Across six fine-grained visual benchmarks at 4B and 9B scales, VAD outperforms direct privileged-view distillation and visual-advantage weighting. Token- level and controlled-target analyses show that the proxy-aligned component is enriched in task-relevant visual corrections and yields stronger target shifts, especially when evidence refutes a mistaken answer. These results support counterfactual target reconstruction as an effective alternative to source-mixed supervision.