When Do We Need On-Policy Distillation? Distilling on Offline Student Rollouts Is Often Better

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Although online distillation is widely adopted, it remains inefficient and does not necessarily outperform offline strategies. This work re-examines the impact of teacher-student alignment on distillation efficacy, revealing that online sampling is only necessary under high token overlap ratios and that long contexts induce signal attenuation. Motivated by these findings, we propose a semi-online distillation method that leverages offline trajectories from an initial student model to enable efficient knowledge transfer. Across 17 comparative experiments involving diverse model configurations, our approach achieves superior performance in 14 cases, yielding accuracy improvements of up to 13.6% while accelerating training speed by 11.4×. These results validate the stability and efficiency advantages of the proposed semi-online framework over conventional distillation paradigms.
📝 Abstract
On-policy distillation (OPD) has become increasingly popular for transferring teacher capabilities to student models. In this work, we ask a critical research question: Is on-policy sampling always beneficial for distilling arbitrary teacher-student pairs? We show that a simple alternative, Semi-OPD, which distills from offline rollouts generated by the initial student, can often outperform OPD in both accuracy and training efficiency. Across 17 teacher-student pairs ranging from 1.5B to 235B parameters, Semi-OPD outperforms OPD in 14 cases, with up to +13.6% accuracy and 11.4x training speedup. We further find that the choice between OPD and Semi-OPD depends on the alignment between the initial teacher and student, quantified by an output-token overlap ratio: OPD is beneficial only when the two are highly aligned with high overlap ratios. Our deeper investigation suggests that effective distillation requires on-policyness w.r.t. both the student and the teacher. For misaligned pairs, student rollouts can become increasingly off-policy w.r.t. the teacher as context length grows, weakening the distillation signal. In contrast, Semi-OPD is often more stable, as it distills on shorter contexts while covering full trajectories and exposing the student to more teacher-preferred tokens. Beyond proposing Semi-OPD as an efficient alternative, our work motivates the community to rethink when to use OPD and to study stronger OPD variants with meaningful teacher-student pairs.
Problem

Research questions and friction points this paper is trying to address.

knowledge distillation
on-policy distillation
teacher-student alignment
language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Distillation
Semi-OPD
Knowledge Distillation
Offline Rollouts
Teacher-Student Alignment
🔎 Similar Papers
2024-07-21arXiv.orgCitations: 1