On the Off-Policy Teacher in On-Policy Distillation

๐Ÿ“… 2026-09-29
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the performance degradation of teacher models in online knowledge distillation caused by processing off-policy prefixes generated by student models. To tackle this issue, we propose SCOUT, a framework that systematically mitigates off-policy bias for the first time. Through co-training, the teacher adapts to student-generated trajectories, while reinforcement learning with verifiable rewards drives adaptive teacher updates to optimize its conditional continuation capability. Our approach significantly enhances the quality of teacher continuations given student prefixes and consistently improves distillation effectiveness across diverse configurations and reasoning tasks.
๐Ÿ“ Abstract
On-policy distillation (OPD) has recently emerged as a promising post-training paradigm in which the student learns from trajectories generated by its own policy under dense teacher supervision. However, OPD introduces a fundamental asymmetry: although the sampled trajectories are on-policy for the student, they are off-policy for the teacher. The teacher is typically optimized to continue from prefixes generated by its own policy, but during OPD it must instead supervise prefixes generated by the student. Empirically, we find that its continuation performance degrades as these prefixes grow longer. To address this issue, we propose Student-COnditioned Updates of the Teacher (SCOUT), a co-training framework that adapts the teacher to student-generated prefixes. Alongside standard OPD updates, SCOUT periodically optimizes the teacher's conditional ability using reinforcement learning with verifiable rewards, where the teacher generates continuations from student prefixes and learns from outcome rewards. Controlled experiments show that SCOUT improves the teacher's ability to continue from student-generated prefixes, supporting the intended mechanism of student-conditioned teacher adaptation. Across multiple teacher--student configurations, model scales, and reasoning domains, SCOUT also consistently improves the effectiveness of on-policy distillation.
Problem

Research questions and friction points this paper is trying to address.

On-Policy Distillation
Off-Policy Teacher
Prefix Degradation
Teacher-Student Asymmetry
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Distillation
Off-Policy Teacher
SCOUT
Reinforcement Learning
Co-training
๐Ÿ”Ž Similar Papers
No similar papers found.