Understanding Off- vs On-Policy Distillation: A Tale of Distinct Training Objectives

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear origins of advantages and vulnerabilities associated with off-policy and on-policy mechanisms in multi-teacher knowledge distillation. By building upon sequential distillation and Kullback-Leibler divergence optimization, this work comparatively analyzes different aggregation objectives. Integrating log-regret bound theory with function approximation, it rigorously investigates the preference-preservation properties, sensitivity to low-probability perturbations, and prefix bias mechanisms inherent in reverse KL divergence. The primary contribution lies in establishing strict theoretical bounds that precisely characterize how feedback sensitivity and long-range dependencies influence distillation performance. Ultimately, this research elucidates the underlying causes of performance discrepancies and identifies critical stability bottlenecks within multi-teacher knowledge distillation frameworks.
📝 Abstract
On-policy distillation (OPD) learns from teacher feedback on student-generated responses and has shown promise in reducing forgetting relative to supervised fine-tuning (SFT). However, its benefits and fragility remain incompletely understood. We study sequential distillation from multiple teachers, where the student minimizes its average divergence from the teachers. Forward Kullback--Leibler (KL) divergence yields a weighted arithmetic mixture, while reverse KL yields a normalized weighted geometric aggregate. We develop algorithms that learn these targets under off-policy and on-policy feedback, respectively, establishing logarithmic regret bounds in the tabular setting and extending the analysis to function approximation. By analyzing these aggregation targets, we identify mechanisms that help explain both the benefits and fragility of OPD. Relative to forward KL, reverse KL can better retain a confident expert's preferences under uninformative feedback, but is more sensitive to teachers that assign very low probabilities to correct responses. Its token-level conditionals also reveal a dependence on continuation distributions that can favor incorrect prefixes over long horizons.
Problem

Research questions and friction points this paper is trying to address.

Knowledge Distillation
On-policy Distillation
Off-policy Distillation
Training Objectives
KL Divergence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Knowledge Distillation
On-Policy vs Off-Policy
KL Divergence
Regret Bounds
Sequential Distillation
🔎 Similar Papers
2024-07-21arXiv.orgCitations: 1
💼 Related Jobs
No related jobs found.
Q
Qiwei Di
Department of Computer Science, University of California, Los Angeles, CA 90095, USA
X
Xuheng Li
Department of Computer Science, University of California, Los Angeles, CA 90095, USA
Kaixuan Ji
Kaixuan Ji
University of California, Los Angeles
machine learning
C
Chenggong Zhang
Department of Computer Science, University of California, Los Angeles, CA 90095, USA
Heyang Zhao
Heyang Zhao
UCLA
Machine Learning
Quanquan Gu
Quanquan Gu
Associate Professor of Computer Science, UCLA
AGILarge Language ModelsReinforcement LearningNonconvex Optimization