🤖 AI Summary
This study addresses the substantial memory overhead of full-vocabulary rectification in online knowledge distillation, where existing sparse methods often introduce noise or bias into rectification semantics. To overcome these limitations, this work proposes the SparseOPD framework, which identifies critical tokens based on rectification magnitude and performs backpropagation exclusively through selected logits. By integrating sign-based residual compensation with rectification-aware dynamic budget allocation, SparseOPD significantly reduces computational costs while preserving rectification quality. Empirical evaluations demonstrate that the proposed method outperforms sampling and Top-K baselines across six task categories, achieving performance comparable to full-vocabulary approaches. Notably, on mathematical reasoning tasks, SparseOPD attains a 99% gradient cosine similarity relative to full-vocabulary rectification while reducing backpropagation memory consumption by 70.5%.
📝 Abstract
On-policy distillation (OPD) uses teacher correction on student-generated responses. Full-vocabulary correction can provide important corrections even for tokens that the student assigns low probability, but backpropagating through all token logits becomes memory-intensive for long sequences. Existing memory-saving approaches estimate corrections from sampled tokens or restrict supervision to the student's TopK tokens, introducing sampling noise or changing the full-vocabulary correction. We introduce \textbf{SparseOPD}, which uses full-vocabulary teacher correction to determine which corrections matter before selecting the token logits to differentiate. SparseOPD first constructs the full-vocabulary correction without retaining its backward graph, then selects tokens by correction magnitude rather than student probability. Signed residual compensation preserves the total promoting and suppressing correction mass, while correction-aware budget allocation distributes the sparse support across positions. Finally, the update backpropagates only through the selected token logits. Across six task--scale settings spanning mathematics, chemistry QA, and multimodal reasoning, SparseOPD outperforms Sampled Token and TopK in task-average accuracy and matches or exceeds Full Vocabulary. Gradient cosine similarity reaches 99\% on 4B mathematics, while 8K full-parameter profiling shows 70.5\% lower backward memory.