Look Before You Select: Rethinking Vocabulary Sparsification in On-Policy Distillation

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the substantial memory overhead of full-vocabulary rectification in online knowledge distillation, where existing sparse methods often introduce noise or bias into rectification semantics. To overcome these limitations, this work proposes the SparseOPD framework, which identifies critical tokens based on rectification magnitude and performs backpropagation exclusively through selected logits. By integrating sign-based residual compensation with rectification-aware dynamic budget allocation, SparseOPD significantly reduces computational costs while preserving rectification quality. Empirical evaluations demonstrate that the proposed method outperforms sampling and Top-K baselines across six task categories, achieving performance comparable to full-vocabulary approaches. Notably, on mathematical reasoning tasks, SparseOPD attains a 99% gradient cosine similarity relative to full-vocabulary rectification while reducing backpropagation memory consumption by 70.5%.
📝 Abstract
On-policy distillation (OPD) uses teacher correction on student-generated responses. Full-vocabulary correction can provide important corrections even for tokens that the student assigns low probability, but backpropagating through all token logits becomes memory-intensive for long sequences. Existing memory-saving approaches estimate corrections from sampled tokens or restrict supervision to the student's TopK tokens, introducing sampling noise or changing the full-vocabulary correction. We introduce \textbf{SparseOPD}, which uses full-vocabulary teacher correction to determine which corrections matter before selecting the token logits to differentiate. SparseOPD first constructs the full-vocabulary correction without retaining its backward graph, then selects tokens by correction magnitude rather than student probability. Signed residual compensation preserves the total promoting and suppressing correction mass, while correction-aware budget allocation distributes the sparse support across positions. Finally, the update backpropagates only through the selected token logits. Across six task--scale settings spanning mathematics, chemistry QA, and multimodal reasoning, SparseOPD outperforms Sampled Token and TopK in task-average accuracy and matches or exceeds Full Vocabulary. Gradient cosine similarity reaches 99\% on 4B mathematics, while 8K full-parameter profiling shows 70.5\% lower backward memory.
Problem

Research questions and friction points this paper is trying to address.

On-policy distillation
vocabulary sparsification
memory efficiency
full-vocabulary correction
knowledge distillation
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Distillation
Vocabulary Sparsification
SparseOPD
Signed Residual Compensation
Memory Efficiency
🔎 Similar Papers
No similar papers found.
Y
Yongliang Miao
The Chinese University of Hong Kong, Shenzhen
S
Shuang Liu
Carnegie Mellon University
Y
Yanguang Liu
New Jersey Institute of Technology
Y
Yandong Bai
Kuaishou
Mengnan Du
Mengnan Du
Assistant Professor, New Jersey Institute of Technology
ExplainabilityNatural Language ProcessingTrustworthy AI