Is Better Teacher Supervision Enough? Unlocking Student-side Learning in Multimodal On-Policy Distillation

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the bottleneck in multimodal online policy distillation, where limited student perceptual capacity constrains the performance ceiling achievable through teacher supervision alone. To overcome this limitation, we propose the S-OPD framework, which explicitly enhances student visual understanding via teacher-calibrated strategy contrast and a policy consistency objective. Furthermore, the framework strengthens perceptual learning on the student side by integrating a token-level gating mechanism with policy alignment under image masking and noise perturbation. Notably, S-OPD can be seamlessly incorporated into existing methods without requiring additional data or parameters. Extensive experiments demonstrate significant performance improvements across eight benchmarks, achieving gains of up to 4.25 points on LogicVista. Moreover, the proposed approach yields complementary benefits when combined with teacher-side optimization techniques, establishing it as an effective and versatile solution for enhancing multimodal policy distillation.
📝 Abstract
On-policy distillation (OPD) improves reasoning by providing token-level supervision from a teacher on a student's own trajectories. Existing methods primarily focus on enhancing this teacher-side guidance (e.g., by enriching teacher inputs and refining teacher feedback), yet we find that limited student perception is another critical bottleneck in multimodal OPD. By providing oracle visual facts, the performance of OPD-trained students can still be substantially improved for both weak and strong teachers. To address this bottleneck, we propose S-OPD, a simple multimodal on-policy distillation framework that explicitly strengthens student perceptual learning through two objectives. Specifically, Teacher-calibrated Policy Contrast separates student policies under original and masked images with teacher-based token-level gating, strengthening the student's reliance on visual evidence during reasoning. Policy Agreement aligns student policies under original and noise-perturbed images, further improving perceptual robustness to visual noise. Notably, our method can be seamlessly plugged into existing OPD frameworks, requiring no additional data annotations, model parameters or inference operations. Extensive experiments on eight benchmarks across student scales and distillation paradigms demonstrate consistent performance improvements, with gains of up to 4.25 points on LogicVista. When combined with existing teacher-side supervision methods, our method can yield further gains. Code is available at https://github.com/Sirilaw/S-OPD.
Problem

Research questions and friction points this paper is trying to address.

Multimodal On-Policy Distillation
Student Perception
Knowledge Distillation
Visual Reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Distillation
Multimodal Reasoning
Perceptual Learning
Policy Contrast
Knowledge Distillation
🔎 Similar Papers
No similar papers found.