๐ค AI Summary
This study addresses the performance degradation of teacher models in online knowledge distillation caused by processing off-policy prefixes generated by student models. To tackle this issue, we propose SCOUT, a framework that systematically mitigates off-policy bias for the first time. Through co-training, the teacher adapts to student-generated trajectories, while reinforcement learning with verifiable rewards drives adaptive teacher updates to optimize its conditional continuation capability. Our approach significantly enhances the quality of teacher continuations given student prefixes and consistently improves distillation effectiveness across diverse configurations and reasoning tasks.
๐ Abstract
On-policy distillation (OPD) has recently emerged as a promising post-training paradigm in which the student learns from trajectories generated by its own policy under dense teacher supervision. However, OPD introduces a fundamental asymmetry: although the sampled trajectories are on-policy for the student, they are off-policy for the teacher. The teacher is typically optimized to continue from prefixes generated by its own policy, but during OPD it must instead supervise prefixes generated by the student. Empirically, we find that its continuation performance degrades as these prefixes grow longer. To address this issue, we propose Student-COnditioned Updates of the Teacher (SCOUT), a co-training framework that adapts the teacher to student-generated prefixes. Alongside standard OPD updates, SCOUT periodically optimizes the teacher's conditional ability using reinforcement learning with verifiable rewards, where the teacher generates continuations from student prefixes and learns from outcome rewards. Controlled experiments show that SCOUT improves the teacher's ability to continue from student-generated prefixes, supporting the intended mechanism of student-conditioned teacher adaptation. Across multiple teacher--student configurations, model scales, and reasoning domains, SCOUT also consistently improves the effectiveness of on-policy distillation.