🤖 AI Summary
This study addresses the high computational cost of teacher supervision in online knowledge distillation by proposing an adaptive supervision selection mechanism that leverages the student model's own successful trajectories as a reference. Specifically, this method precisely allocates teacher resources by filtering out failed samples, while integrating hidden state trajectory discrepancy analysis with a budget-aware token selection strategy to achieve efficient supervision. Experimental results demonstrate that the proposed approach maintains comparable inference performance while requiring only 3.46% to 5.02% of the teacher inputs, thereby substantially reducing the computational overhead associated with online distillation.
📝 Abstract
On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher, but providing such supervision for every rollout requires substantial teacher computation. We introduce Success-Referenced On-Policy Distillation (SR-OPD), which reduces this cost by selecting which prompts and rollouts receive teacher supervision. When the student produces both successful and failed rollouts for the same prompt, a successful rollout can serve as a natural reference for selecting failed rollouts. SR-OPD therefore focuses on such prompts and prioritizes failed rollouts whose hidden-state trajectories show sustained divergence from a successful reference, while accounting for estimated teacher-input cost. Across three teacher-student pairs and six mathematical reasoning benchmarks, SR-OPD uses only 3.46-5.02% of the teacher-input tokens required by Vanilla OPD in the one-pass setting while maintaining comparable reasoning performance. Under a controlled setting matched to 5% of Vanilla OPD's teacher-input budget, further experiments support both key design choices: focusing supervision on prompts with both successful and failed rollouts, and using successful rollouts to guide failure selection. These results indicate that a student's own successful behavior can serve as a useful reference for allocating teacher supervision under a fixed teacher-input budget.