PAST: Privileged Adaptation from Complete Student Trajectories for On-Policy Self-Distillation

๐Ÿ“… 2026-08-09
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses a key limitation in existing on-policy self-distillation methods, which fail to effectively leverage privileged knowledge embedded in post-hoc feedback (e.g., success/failure outcomes) from student trajectories. The authors propose PAST, a novel approach that, for the first time, utilizes complete student trajectories as privileged information to adaptively refine the teacher model. While preserving the studentโ€™s distillation prefix, PAST employs trajectory-conditioned distillation to disentangle transferable policy shifts from trajectory-specific variations and theoretically characterizes the teacherโ€™s capacity to convey knowledge to a prefix-only student. The method integrates Forward-KL distillation, student-proximity regularization, and a distribution-preserving mechanism over correct trajectories. Evaluated on three mathematical reasoning benchmarks, PAST achieves a 5.6 percentage point improvement in Avg@12 macro-average over vanilla OPSD, with ablation studies confirming the critical roles of trajectory completeness and teacher adaptivity.
๐Ÿ“ Abstract
On-policy self-distillation (OPSD) uses a privileged teacher to supervise a reasoning model on prefixes sampled from its own rollouts. Yet each rollout also reveals how the student's response unfolds and whether it succeeds, student-specific hindsight that standard OPSD does not use to form the teacher. We introduce Privileged Adaptation from Student Trajectories (PAST), which treats each completed student trajectory as additional privileged information for the OPSD teacher while leaving the student's distillation prefixes unchanged. PAST preserves the student's next-token distribution on correct trajectories and uses failed trajectories to adapt the teacher toward verified success under student-proximity regularization. We characterize what such a trajectory-conditioned teacher can transfer to a prefix-only student. Forward-KL distillation projects the teacher distributions to their conditional arithmetic mean given the prefix. This projection separates trajectory-specific variation that remains privileged from the mean policy shift available to the student. For correct trajectories, the unclipped population objective also has the frozen student as an ideal distributional fixed point. Across three mathematical reasoning benchmarks, PAST improves the Avg@12 macro average over Vanilla OPSD by 5.6 percentage points. A $2\times2$ factorial study shows gains from both complete-trajectory access and teacher adaptation, while trajectory removal and shuffling confirm that the adapted teacher uses the matching hindsight context.
Problem

Research questions and friction points this paper is trying to address.

on-policy self-distillation
privileged information
student trajectories
reasoning models
teacher adaptation
Innovation

Methods, ideas, or system contributions that make the work stand out.

on-policy self-distillation
privileged adaptation
student trajectories
trajectory-conditioned teacher
forward-KL distillation
๐Ÿ”Ž Similar Papers
2024-07-21arXiv.orgCitations: 1
๐Ÿ’ผ Related Jobs
No related jobs found.
Y
Yangyang Feng
The Hong Kong University of Science and Technology (Guangzhou)
Z
Zhuoyan Feng
Sun Yat-sen University
J
Junlan Chen
The Hong Kong University of Science and Technology (Guangzhou)