Flash-OPD: Fast On-Policy Distillation

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of fixed trajectory lengths in online distillation, which often leads to truncated supervision or computational redundancy. To overcome this, we propose an adaptive trajectory generation framework grounded in reliability bounds. This framework shifts the control paradigm from rolling-horizon planning to trajectory-level bound verification, thereby decoupling event scheduling from precise stopping decisions. By interleaving cached generation with teacher verification and employing a cumulative compatibility counting mechanism, it dynamically determines when to terminate trajectory generation. Consequently, the proposed approach achieves a 2.2× to 7.5× training speedup over standard methods while preserving model accuracy, significantly enhancing the efficiency of online distillation.
📝 Abstract
On-policy distillation (OPD) provides dense teacher supervision on student-generated trajectories, but generating and evaluating long rollouts incurs substantial training cost. Existing acceleration methods reduce this cost through open-loop rollout schedules or closed-loop horizon adaptation. However, supervision compatibility can vary substantially across trajectories, making a single rollout horizon difficult to match their heterogeneous reliable lengths: an overly short horizon may truncate useful supervision, while an overly long one wastes computation beyond reliable regions. Our key insight is that the trajectory-specific reliability boundary need not be predicted before generation. By viewing reliability as the first-passage of accumulated low teacher--student compatibility events, the boundary is inherently unknown before sampling, yet whether it has been reached can be determined exactly from the observed prefix. Building on this insight, we propose *Flash-OPD*, which shifts from rollout-horizon control to adaptive trajectory-level boundary verification. *Flash-OPD* interleaves cached generation with teacher verification and independently stops each trajectory according to its observed compatibility events. To reduce verification overhead, the recent event rate is used only to schedule the next verification point, while the actual stopping decision always relies on the exact cumulative count. This separation prevents estimation errors from causing premature termination while enabling efficient verification during generation. Extensive experiments across diverse datasets and teacher--student settings show that *Flash-OPD* achieves $2.2\times$--$7.5\times$ speedups over standard OPD while maintaining or improving accuracy.
Problem

Research questions and friction points this paper is trying to address.

on-policy distillation
rollout horizon
training efficiency
trajectory reliability
knowledge distillation
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Distillation
Adaptive Trajectory Boundary
First-Passage Reliability
Interleaved Verification
Knowledge Distillation
🔎 Similar Papers