On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge in existing knowledge distillation research where multivariate confounding obscures the independent effects of policy selection on forgetting, sparsity, and generalization. Leveraging Llama3 and Qwen2.5 models, we construct a controlled strong-to-weak distillation framework that employs token-level KL divergence optimization and gradient analysis to independently decouple policies, KL directions, and learning rates across a continuous policy spectrum. Our findings reveal that forward KL exhibits policy robustness whereas reverse KL is highly sensitive, and that learning rate predominantly governs forgetting phenomena. Furthermore, this work challenges the presumed inherent superiority of on-policy data, demonstrating that its utility depends on objective and hyperparameter configurations, and that its generalization advantages in subsequent RLVR remain unreliable. These conclusions exhibit broad robustness across multiple models and reasoning tasks.
📝 Abstract
On-policy learning has been argued to reduce catastrophic forgetting, produce sparser parameter updates, and improve generalisation. However, existing comparisons between supervised fine-tuning and reinforcement learning vary many factors simultaneously, making the contribution of rollout policy difficult to isolate. We study the effect of rollout policy in a controlled strong-to-weak distillation setting, by independently varying rollout policy, token-level KL direction, and learning rate across the Llama3 and Qwen2.5 model families and reasoning tasks spanning scientific, medical, and arithmetic domains. Our analysis reveals a nuanced picture of distillation dynamics in which rollout policy does not necessarily play a central role. Instead, token-level KL direction more clearly shapes task performance and output coverage, while learning rate governs forgetting and update sparsity. Analysis of KL gradients and experiments along a continuous student-teacher rollout-policy spectrum explain this pattern: forward KL is remarkably robust to rollout policy, with its performance stable and strong despite changes to the rollout policy, whereas reverse KL is substantially more sensitive and favours student-generated rollouts. On-policy data nevertheless improves generalisation to harder variants of the Countdown arithmetic task under both KL directions, although this advantage does not reliably persist after subsequent RLVR. Our broader conclusions remain robust to removing gradient clipping, using sampled KL estimators, and training on tasks requiring longer reasoning chains. Overall, our results challenge the view that on-policy rollouts are inherently preferable and show that their value depends critically on the objective, evaluation setting, and optimisation hyperparameters.
Problem

Research questions and friction points this paper is trying to address.

knowledge distillation
on-policy learning
off-policy learning
rollout policy
KL divergence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Knowledge Distillation
On-Policy Learning
KL Divergence
Rollout Policy
Catastrophic Forgetting
🔎 Similar Papers
2024-07-21arXiv.orgCitations: 1