Reward-Aligned Reweighting for On-Policy Distillation

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of uniform weighting in online distillation, which overlooks how decisions impact subsequent reasoning and may suppress viable strategies or reinforce unreliable paths. We propose R²-OPD, a method that dynamically reallocates teacher supervision weights based on outcome consistency across verification trajectories and the magnitude of student-teacher divergence, thereby amplifying reward-aligned corrections. Theoretically establishing conditions under which reallocation enhances task progress, our approach surpasses both uniform imitation and hard filtering to enable precise guidance under dense feedback. Empirically, R²-OPD achieves the highest average accuracy across seven mathematical benchmarks, outperforming standard online process distillation by 3.5 and 2.4 percentage points, respectively, while also demonstrating significant improvements over baselines in code generation tasks.
📝 Abstract
On-policy distillation (OPD) trains a student language model with dense feedback from a stronger teacher on student-generated trajectories. Yet standard OPD weights token-level distillation terms uniformly, implicitly treating local teacher preference as a proxy for correction utility. A decision's task value, however, depends on how the student completes the subsequent reasoning. This mismatch can cause imitation to suppress viable student strategies or reinforce paths the student cannot reliably execute. Verified trajectory outcomes provide complementary evidence about continuation quality, but do not directly identify the utility of individual decisions. We introduce Reward-Aligned Reweighting for On-Policy Distillation (R$^{2}$-OPD), which uses outcome agreement and the magnitude of teacher--student disagreement to continuously reallocate teacher supervision. It gives reward-aligned corrections greater relative influence while retaining dense feedback, moving beyond uniform imitation and hard filtering. Our analysis formalizes the mismatch between local teacher preference and student continuation value and establishes sufficient conditions for reallocation to improve first-order task progress over uniform OPD. Across seven mathematical reasoning benchmarks, R$^{2}$-OPD achieves the highest average accuracy among the compared training methods in both cross-size and same-size distillation. It outperforms standard OPD on all seven benchmarks, with average gains of 3.5 and 2.4 percentage points for 1.7B and 4B students, respectively. An extension to code generation yields an average gain of 1.6 percentage points over standard OPD. These results highlight outcome-guided supervision allocation as an effective way to translate dense teacher feedback into stronger student performance across model scales and task domains.
Problem

Research questions and friction points this paper is trying to address.

On-Policy Distillation
Reward Alignment
Knowledge Distillation
Language Models
Reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Distillation
Reward-Aligned Reweighting
Outcome-Guided Supervision
Teacher-Student Disagreement
Knowledge Distillation