Dr. OPD: Learning What to Follow for Optimal On-Policy Distillation of Large Language Models

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of conventional online knowledge distillation, which treats all teacher signals equally while overlooking the differential impact of individual tokens on model performance. We propose an optimal weighted online distillation framework that formulates the distillation process as a bilevel optimization problem. By employing an iterative solver combined with gradient descent, our method dynamically adjusts token-level weights, and we further derive a closed-form weight update solution to maximize the expected reward of the student model. Theoretical analysis rigorously demonstrates that this strategy strictly outperforms uniform weighting baselines. Empirical evaluations on mathematical reasoning and code generation tasks show that our approach significantly surpasses existing baselines, achieving an average improvement of 9.7 points in strong-to-weak distillation scenarios and enabling smaller student models to exceed their teachers' performance.
📝 Abstract
On-policy distillation (OPD) trains a student on its own generated responses using dense, token-level supervision from a stronger teacher. Vanilla OPD treats all teacher signals equally, assuming that the teacher's supervision is equally important for every token. However, teacher signals at different tokens may have very different effects on the student's performance: some correct important reasoning errors, while others have little effect on the final answer. Motivated by this observation, we introduce Dr. OPD (OPD Done Right), which defines the optimal weighted OPD to maximize the student's performance. We formulate Dr. OPD as a bilevel optimization problem in which the student learns from weighted teacher supervision, while the weights are selected to maximize the expected reward of the resulting student. To solve Dr. OPD, we develop an efficient iterative solver that updates the token weights and student policy alternatively. At each round, it updates weights in closed form and then takes one gradient step on the resulting weighted OPD objective. Under regularity conditions, we show that this weighted update achieves a higher expected reward than a vanilla OPD update. Empirically, across strong-to-weak and same-size distillation on math and code, Dr. OPD consistently outperforms all evaluated baselines. In particular, in the strong-to-weak distillation setting, Dr. OPD improves average math performance by $9.7$ points over vanilla OPD, and enables the smaller student to surpass its larger teacher.
Problem

Research questions and friction points this paper is trying to address.

On-policy distillation
Large language models
Token-level supervision
Knowledge distillation
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-policy distillation
Bilevel optimization
Token-level weighting
Iterative solver
Large language models
🔎 Similar Papers
No similar papers found.