Teaching a Moving Student: Rethinking the Curriculum of On-Policy Distillation

📅 2026-09-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge in online policy distillation where the continuous evolution of student policies causes early trajectory divergences to be overlooked, rendering fixed sampling or replay strategies suboptimal for supervision budget utilization. To overcome this, we propose R-OPD, a curriculum learning framework that introduces a novel adaptive scheduling strategy based on a gradient-triggering mechanism. By performing minibatch-level variance detection to monitor gradient dynamics, our method precisely identifies the onset of policy degradation and dynamically switches to initial policy replay, thereby facilitating efficient knowledge transfer. Extensive experiments demonstrate that R-OPD significantly improves accuracy across multiple mathematical reasoning benchmarks, achieving gains of up to 6.15%, while generating longer reasoning chains that fully unlock the model's potential.
📝 Abstract
In on-policy distillation (OPD), the student determines which states receive teacher supervision. As its policy evolves, earlier response prefixes become less likely even though teacher-student disagreement on them persists. Under matched trajectory and optimization budgets, neither more queries nor more frequent rollout resampling is uniformly beneficial. Current-policy rollouts outperform initial-policy rollouts at shorter response budgets, but this ranking reverses at longer budgets. Replaying initial-policy rollouts after current-policy training improves accuracy, whereas replaying fixed recent rollouts does not reproduce the gain. We propose R-OPD, a gradient-triggered curriculum that adaptively selects when to revisit initial-student trajectories. When changes in mean gradients fall within minibatch-level variation for two consecutive comparisons, training switches from the next iteration onward to initial-policy replay. Across eight mathematics benchmarks and three training-data orders, R-OPD improves average accuracy over continued current-policy sampling by 2.25/4.44 percentage points at 16K/32K for a 0.6B student and 2.81/6.15 points for a 1.7B student. With a 30B-A3B teacher, it improves 8B accuracy by 4.05 points at 32K. At 32K, R-OPD exceeds a fixed schedule of 40 current-policy updates followed by 20 replay updates by 1.39/1.66/1.85 points for 0.6B/1.7B/8B. At the same generation cap, R-OPD also produces longer responses, suggesting that well-timed revisits help students use more of their reasoning capacity.
Problem

Research questions and friction points this paper is trying to address.

On-Policy Distillation
Knowledge Distillation
Curriculum Learning
Policy Evolution
Trajectory Replay
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Distillation
Curriculum Learning
Gradient-triggered Replay
Knowledge Distillation
Adaptive Scheduling
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
L
Lingxiang Hu
Tencent
T
Tianle Xia
Tencent
Yiding Sun
Yiding Sun
Renmin University of China
Large Language ModelsExplainable Recommendation
M
Ming Xu
Tencent
L
Linfang Shang
Tencent
L
Lan Xu
Tencent
N
Ning Zheng
Tencent
W
Wei Xu
Tencent
J
Jie Jiang
Tencent