On the Position Bias of On-Policy Distillation

📅 Unknown Date
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses a key limitation of On-Policy Distillation (OPD), where uniform weighting of all tokens introduces low-quality supervision due to increasing distributional divergence between student and teacher models in later training stages, thereby impairing learning efficiency. To mitigate this issue, the authors propose Importance-Weighted OPD (IW-OPD), which adopts a constrained optimization perspective and dynamically adjusts token-level weights based on the cumulative KL divergence between student and teacher distributions. This adaptive reweighting scheme emphasizes reliable early-stage signals while suppressing noisy later-stage supervision. Empirical results demonstrate that IW-OPD significantly accelerates convergence and consistently outperforms standard OPD under both same-scale and cross-scale distillation settings, achieving a performance gain of up to 6.9 points on the AIME-2025 benchmark.
📝 Abstract
On-Policy Distillation (OPD) improves the learning efficiency of standard reinforcement learning through dense, token-level supervision from teachers. In the standard KL objective of OPD, token-level losses are uniformly averaged, implying equal weights for all tokens. However, we discover that not all tokens are created equal: as student rollouts grow longer, they deviate further from the teacher's distribution, leading to degraded supervision quality at later positions. As a result, OPD using only the first 30% of tokens can perform comparably to using all tokens, whereas OPD using only the last 30% of tokens barely learns anything. In this work, we provide a principled understanding of this issue through the lens of constrained optimization. Based on these insights, we derive Importance-Weighted On-Policy Distillation (IW-OPD), in which the weight assigned to each token depends on the accumulated discrepancy between the student's and teacher's distributions, naturally upweighting earlier tokens and downweighting later ones with larger deviations. We show that IW-OPD converges significantly faster than OPD, with better learning efficiency, and achieves better final performance than standard OPD in both same-size and cross-scale settings, improving performance up to 6.9 points on AIME-2025.
Problem

Research questions and friction points this paper is trying to address.

Position Bias
On-Policy Distillation
Token-level Supervision
Distribution Discrepancy
Reinforcement Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Distillation
Position Bias
Importance Weighting
Reinforcement Learning
Constrained Optimization