ReTaCo: Residual-Target Control for On-Policy Distillation

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high transmission cost of full-vocabulary distribution transfer and teacher non-stationarity in online knowledge distillation by proposing ReTaCo. The method reduces communication overhead through top-k probability truncation and residual sign aggregation, while employing a single-sample estimator to optimize the reverse KL divergence. Furthermore, it introduces a residual target control mechanism, theoretically proven to possess a unique optimal solution and monotonic controllability, effectively overcoming the vanishing gradient deficiency inherent in entropy-aware online distillation (EOPD). Experiments demonstrate that ReTaCo significantly outperforms EOPD across multiple mathematical and code reasoning benchmarks, achieving low-cost and highly stable online distillation.
📝 Abstract
On-policy distillation (OPD) trains a student on its own generated prefixes with token-level teacher feedback, but transmitting or storing the teacher's full-vocabulary distribution at every token is costly. Entropy-aware OPD (EOPD) adds forward supervision to reverse KL to help the student recover plausible tokens it underestimates, using only the teacher's top-$k$ probabilities to limit cost. Because EOPD renormalizes these probabilities, its target assigns no mass to the omitted vocabulary. We prove that the resulting loss keeps pushing the student's top-$k$ mass toward one even after the student matches the teacher's relative probabilities within the top-$k$ set, so the teacher itself is not a stationary point whenever the omitted tokens have positive teacher probability. We propose ReTaCo (Residual-Target Control), which keeps the top-$k$ tokens individually and groups the remaining tokens into one residual symbol, and pairs this forward target with a single-sample estimator whose expectation equals the full-vocabulary reverse KL. With teacher top-$k$ mass $m$, the residual target is $(1-β)(1-m)$ for $β\in[0,1]$: $β=0$ preserves the teacher's mass, and larger $β$ moves more mass onto the top-$k$ tokens without changing their relative probabilities. At a fixed prefix, we prove that the population objective has a unique optimum whose top-$k$ mass lies between $m$ and $m+β(1-m)$ and increases monotonically with $β$; at $β=0$, underestimated top-$k$ tokens still receive non-vanishing recovery gradients. Numerical optimization confirms these predictions, and across three teacher-student pairs, ReTaCo outperforms EOPD on most mathematics and code benchmarks.
Problem

Research questions and friction points this paper is trying to address.

on-policy distillation
entropy-aware OPD
top-k approximation
renormalization bias
knowledge distillation
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-policy distillation
Residual-target control
Reverse KL divergence
Knowledge distillation
Top-k approximation
Z
Zixiang Ni
Xi’an Jiaotong University
Z
Zhuo Hu
Zhejiang University
R
Renjie Cao
Georgia Institute of Technology
W
Weijie Ren
Zhejiang University
B
Binqin Shi
Xi’an Jiaotong University
W
Weijia Zhang
Yale University
S
Shuheng Cao
University of California, San Diego
Z
Zhicheng Shi
Zhejiang University
Z
Zhenhao Zhang
Tsinghua University
Haomin Wen
Haomin Wen
Carnegie Mellon University
Data MiningUrban ComputingSpatio-Temporal Data MiningFoundation Model
Zhiyuan Hu
Zhiyuan Hu
National University of Singapore, Massachusetts Institute of Technology
Natural Language Processing