🤖 AI Summary
This study addresses the limitation of teacher supervision quality in self-distillation by proposing the B-OPSD framework. This method innovatively recycles optimization trajectories to construct a "future teacher": the model is first further trained to yield a stronger supervisor, then rolled back to its student state to receive superior guidance from this future teacher. Technically, B-OPSD integrates online policy self-distillation, privileged-information trajectory generation, and token-level fine-grained supervision to enable backward knowledge distillation and self-improvement. Evaluated on mathematical reasoning tasks with Qwen3, B-OPSD significantly outperforms standard methods, achieving substantial accuracy improvements. These results validate the effectiveness of leveraging optimization progress as a high-quality supervisory signal for enhancing model capabilities through self-distillation.
📝 Abstract
On-policy self-distillation (OPSD) improves large language models by letting a self-teacher with privileged information provide dense token-level supervision on the model's own trajectories. Yet existing methods typically construct the self-teacher from the current, initial, or slowly averaged policy state, leaving the quality of supervision constrained by the teacher's ability to exploit privileged information. We ask whether the model's own optimization progress can instead be recycled into a stronger self-teacher. In this paper, we introduce Bootstrapped On-Policy Self-Distillation (B-OPSD), which temporarily trains the policy ahead to obtain a future teacher, restores the student to the original policy state, and then uses the future teacher to supervise the restarted student. The future teacher improves supervision in two complementary ways, it can generate more reliable privileged trajectories and, conditioned on them, provide more informative token-level targets along the restarted student's on-policy trajectories. Experiments on mathematical reasoning with Qwen3-4B and Qwen3-8B show consistent improvements over standard OPSD in both settings, including gains from 27.50 to 41.30 and from 48.80 to 64.44 in the rollout-privileged setting. Our findings point to a broader principle for self-improving models that future learning progress can be distilled backward, preserving acquired knowledge while bootstrapping beyond the optimization state that produced it.