ResOPD: Tail Residualization for Sparse On-Policy Distillation

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inherent trade-off in sparse online distillation between the high variance of sampling-based estimation and the bias introduced by Top-k optimization. To overcome this challenge, we propose ResOPD, whose core innovation lies in a tail residualization mechanism. Specifically, unobserved vocabulary tokens are aggregated into a coarse tail event, enabling sampling exclusively over fine-grained residuals to achieve unbiased full-vocabulary gradient estimation without additional teacher queries. By integrating an unbiased reverse KL estimator with a sparse communication interface, ResOPD substantially reduces gradient variance and stabilizes online training dynamics. Experimental results demonstrate that ResOPD significantly improves downstream task performance, establishing it as an efficient, plug-and-play primitive suitable for broad deployment.
📝 Abstract
On-policy distillation (OPD) trains a student model on self-generated trajectories, but transmitting dense teacher distributions across long reasoning traces creates prohibitive communication and memory bottlenecks. Practical systems therefore rely on sparse teacher interfaces, typically transmitting either the sampled-token score or a small Top-$k$ distribution. However, this sparse setting faces a fundamental dilemma: sampled-token estimators are unbiased but suffer from severe gradient variance, whereas directly optimizing Top-$k$ objectives introduces systematic bias. To improve this trade-off, we propose ResOPD (On-Policy Distillation with Tail Residualization), which provides unbiased full-vocabulary reverse KL gradient estimation under on-policy sampling, with substantial variance reduction under the same sparse payload in the evaluated settings. ResOPD aggregates the unobserved vocabulary into an observable coarse tail event, computes its exact aggregate gradient, and samples only the fine-grained within-tail residual, which requires no additional teacher queries or forward passes. Extensive experiments demonstrate that ResOPD substantially reduces gradient variance, stabilizes online training dynamics, and improves downstream performance across the evaluated settings. These results establish ResOPD as an efficient, plug-and-play variance reduction primitive for sparse on-policy distillation.
Problem

Research questions and friction points this paper is trying to address.

On-Policy Distillation
Sparse Teacher Interface
Gradient Variance
Knowledge Distillation
Top-k Distribution
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Distillation
Tail Residualization
Variance Reduction
Sparse Teacher Interface
Reverse KL Divergence
🔎 Similar Papers