TISD: On-Policy Self-Distillation with Trajectory Intervention

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the data collection bottleneck and the absence of downstream contextual supervision in online self-distillation, which arise from evaluating only student-sampled trajectories. To overcome these limitations, we propose Trajectory Intervention Self-Distillation, a method that reframes teacher-student divergence as global trajectory branching proposals rather than local corrections. Specifically, the teacher is compelled to select a branching action, after which the student generates the suffix. A branch-regenerate-distill mechanism then performs privileged contextual distillation over complete trajectories, effectively exposing critical downstream supervision signals. Experimental results demonstrate that the proposed algorithm outperforms SDPO by 1.2 percentage points on Avg@4 for code models and by 0.8 percentage points on Avg@128 under equivalent step budgets in scientific domains.
📝 Abstract
On-policy self-distillation (OPSD) provides dense teacher targets, but evaluates them only along student-sampled rollouts. When the privileged teacher favors an alternative action at a visited prefix, OPSD can provide a target for the branch decision but cannot supervise the successor contexts induced by that action unless the student samples it. This creates a training-time data-collection bottleneck and suggests a different role for teacher-student disagreement: proposing a trajectory branch rather than identifying a sufficient local repair. Our diagnostic framework using controlled token interventions reveals that a teacher-preferred token at peak disagreement can improve student continuation success, while its local corrective value is limited. Motivated by this finding, we introduce a simple branch-regenerate-distill algorithm, Trajectory-Intervention Self-Distillation (TISD). TISD forces a teacher-selected branch action, returns suffix generation to the student, and distills the full trajectory under the privileged-context-conditioned teacher. Across the coding models, TISD improves average Avg@4 over SDPO by 1.2 percentage points. Across the science domains, it improves average Avg@128 by 0.8 points under an equal-step budget and by 0.3 points under an equal-time budget. These results support teacher-guided branching as a way to expose useful successor contexts for self-distillation.
Problem

Research questions and friction points this paper is trying to address.

on-policy self-distillation
trajectory intervention
data-collection bottleneck
teacher-student disagreement
successor contexts
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Self-Distillation
Trajectory Intervention
Teacher-Guided Branching
Token Intervention
Branch-Regenerate-Distill
🔎 Similar Papers
2024-07-21arXiv.orgCitations: 1