Train Ahead, Distill Back: Bootstrapping On-Policy Self-Distillation for Large Language Models

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of teacher supervision quality in self-distillation by proposing the B-OPSD framework. This method innovatively recycles optimization trajectories to construct a "future teacher": the model is first further trained to yield a stronger supervisor, then rolled back to its student state to receive superior guidance from this future teacher. Technically, B-OPSD integrates online policy self-distillation, privileged-information trajectory generation, and token-level fine-grained supervision to enable backward knowledge distillation and self-improvement. Evaluated on mathematical reasoning tasks with Qwen3, B-OPSD significantly outperforms standard methods, achieving substantial accuracy improvements. These results validate the effectiveness of leveraging optimization progress as a high-quality supervisory signal for enhancing model capabilities through self-distillation.
📝 Abstract
On-policy self-distillation (OPSD) improves large language models by letting a self-teacher with privileged information provide dense token-level supervision on the model's own trajectories. Yet existing methods typically construct the self-teacher from the current, initial, or slowly averaged policy state, leaving the quality of supervision constrained by the teacher's ability to exploit privileged information. We ask whether the model's own optimization progress can instead be recycled into a stronger self-teacher. In this paper, we introduce Bootstrapped On-Policy Self-Distillation (B-OPSD), which temporarily trains the policy ahead to obtain a future teacher, restores the student to the original policy state, and then uses the future teacher to supervise the restarted student. The future teacher improves supervision in two complementary ways, it can generate more reliable privileged trajectories and, conditioned on them, provide more informative token-level targets along the restarted student's on-policy trajectories. Experiments on mathematical reasoning with Qwen3-4B and Qwen3-8B show consistent improvements over standard OPSD in both settings, including gains from 27.50 to 41.30 and from 48.80 to 64.44 in the rollout-privileged setting. Our findings point to a broader principle for self-improving models that future learning progress can be distilled backward, preserving acquired knowledge while bootstrapping beyond the optimization state that produced it.
Problem

Research questions and friction points this paper is trying to address.

On-policy self-distillation
Large language models
Self-teacher
Supervision quality
Privileged information
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bootstrapped On-Policy Self-Distillation
Self-Distillation
Large Language Models
Future Teacher
Mathematical Reasoning
🔎 Similar Papers
No similar papers found.
Z
Zheng Zhang
School of Information Science and Technology, ShanghaiTech University; State Key Laboratory of General Artificial Intelligence, BIGAI
X
Xinyue Tan
School of Information Science and Technology, ShanghaiTech University
L
Lufei Li
School of Information Science and Technology, ShanghaiTech University
X
Xinyi Zhang
Hefei University of Technology
Yexin Li
Yexin Li
State Key Laboratory of General Artificial Intelligence BIGAI
reinforcement learningmulti-agent systemmulti-armed banditsdata mining
Kan Ren
Kan Ren
Assistant Professor, ShanghaiTech University
Machine LearningData MiningLarge Language ModelFoundation Model