🤖 AI Summary
This study addresses the high generation costs in online distillation caused by frequent refreshing of student trajectories. To this end, it proposes the REVO framework, which enables repeated optimization without regenerating full trajectories by reusing student trajectories for multi-step updates. Specifically, REVO introduces variance-guided optimization of critical token positions, stable prefix weighting, and a one-step resampling mechanism. Its core contribution lies in integrating off-policy distillation with log-probability ratio analysis to substantially enhance trajectory reuse efficiency. Experimental results demonstrate that REVO matches or surpasses the performance of conventional baselines requiring 200 iterations within only 50 iterations, thereby significantly reducing computational overhead.
📝 Abstract
On-policy distillation (OPD) trains language models using dense token-level teacher supervision on student-generated trajectories. However, its reliance on frequently refreshed student rollouts often incurs substantial generation cost. We introduce REVO, an off-policy distillation framework that improves rollout efficiency by reusing each student rollout for multi-step learner updates. REVO addresses prefix-level and current-token policy mismatch through stabilized prefix weighting and one-step resampling from the current student, which enables repeated updates without regenerating full trajectories. To prioritize informative token positions within reused rollouts, REVO uses the variance of the student-teacher log-probability ratio to quantify the remaining token-level learning signal and guide repeated optimization. Across multiple student-teacher scales, REVO with only 50 rollout iterations matches or exceeds OPD baselines trained for 200 iterations on both in-domain and cross-domain reasoning benchmarks.