Evidence-Aligned Multimodal On-Policy Self-Distillation for Fine-Grained Visual Understanding

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issue in fine-grained visual understanding where teacher models introduce perturbations through lagging or cropping, causing supervision signals to deviate from authentic visual evidence. To mitigate this, we propose an Evidence-Aligned Distillation (EAD) framework that introduces a novel independent evidence reference mechanism by predicting changes on masked original images. By integrating multimodal online policy self-distillation with discrepancy analysis of the current student state and applying cosine similarity weighting, our approach precisely filters out noise corrections not driven by visual evidence. Remarkably, retaining only 6% of high-quality supervision signals enables EAD to consistently outperform existing state-of-the-art methods, significantly enhancing performance in fine-grained visual understanding tasks.
📝 Abstract
Fine-grained visual understanding requires models to recognize small details within complex images. Multimodal on-policy self-distillation (OPSD) addresses this challenge by using a teacher conditioned on evidence-centered crops to supervise a student conditioned on original images along student-generated trajectories. Ideally, teacher corrections, the distributional changes from the student toward the privileged teacher, should be driven by task-relevant visual evidence. However, the designs that make the teacher effective also introduce other interference. Using a lagged or frozen teacher improves training stability but introduces a model-state gap from the evolving student, while cropping enhances task-relevant evidence but also loses the visual context. These two sources of interference make the teacher corrections not purely rely on the visual evidence. We introduce Evidence-Aligned multimodal on-policy self-Distillation (EAD), which retains the crop-conditioned teacher as the target but constructs a separate evidence reference for weighting the corrections. To exclude the effect of lagged model-state from this reference, EAD measures prediction changes using the current student. To avoid crop-induced context changes, EAD masks the evidence region in the original image while preserving the other visual context. The change from the student's masked-image prediction to its original-image prediction provides a controlled reference for the direction in which the visual evidence shifts the student's prediction. EAD weights each teacher correction by its cosine alignment with the reference, i.e., retaining aligned corrections and downweighting the rest. Retaining only 6\% of the supervision mass of dense OPSD, EAD consistently outperforms previous state-of-the-art methods.
Problem

Research questions and friction points this paper is trying to address.

fine-grained visual understanding
multimodal on-policy self-distillation
teacher-student gap
visual context loss
evidence alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Evidence-Aligned Distillation
On-Policy Self-Distillation
Fine-Grained Visual Understanding
Multimodal Learning
Correction Weighting
🔎 Similar Papers
2024-02-20International Conference on Machine LearningCitations: 30