SAPD: Step-Aligned Privileged Distillation

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the high computational cost of online post-training for large language models and the insufficient supervision provided by fixed demonstrations. We propose a rollout-free self-distillation method that departs from the conventional paradigm of treating reference solutions as undifferentiated context. Instead, it innovatively converts demonstrations into context-dependent distributional supervision signals via a step-alignment mechanism. By integrating progress-correlated reasoning transitions with privileged guidance, we construct an efficient offline self-training framework. Experiments demonstrate that our approach significantly outperforms supervised fine-tuning on mathematical reasoning benchmarks while achieving performance comparable to online reinforcement learning. Furthermore, it preserves code generation capabilities and yields approximately a 2× training speedup.
📝 Abstract
On-policy post-training can improve large language models by learning from their own trajectories, but requires costly rollout generation. We ask whether fixed demonstrations can support competitive off-policy learning through better supervision. Our premise is that their usefulness depends not only on the training trajectories, but also on whether supervision provides informative preferences among continuations and connects this guidance to the reasoning decision being learned. We introduce Step-Aligned Privileged Distillation (SAPD), a rollout-free self-distillation method that turns demonstrations into step-aligned distributional supervision. Its key insight is to use the known progression of a reference solution to associate each reasoning transition with targeted privileged guidance, rather than treating the solution as undifferentiated context. On mathematical reasoning benchmarks, SAPD outperforms supervised fine-tuning and label smoothing on average while remaining competitive with on-policy reinforcement learning and self-distillation. Analyses support both the value of context-dependent distributional guidance and the benefit of aligning privileged information with the current step. SAPD also largely preserves out-of-domain coding performance and achieves approximately 2x training-loop speedups over the on-policy baselines. These findings suggest that carefully constructed supervision can make fully off-policy post-training a competitive and computationally efficient alternative. Our code is available at https://github.com/Miaow-Lab/SAPD.
Problem

Research questions and friction points this paper is trying to address.

off-policy learning
post-training
large language models
rollout-free
knowledge distillation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Step-Aligned Privileged Distillation
Rollout-free Self-distillation
Off-policy Post-training
Distributional Supervision
Mathematical Reasoning
🔎 Similar Papers
No similar papers found.
Tianle Wang
Tianle Wang
Brookhaven National Lab
High performance computation
Jiayu Liu
Jiayu Liu
University of Science and Technology of China
Artificial IntelligenceKnowledge LearningMathematical ReasoningNatural Language Processing
R
Ruizhi Zhao
Hong Kong Institute of AI for Science, City University of Hong Kong; Shenzhen University
N
Ning Miao
Department of Data Science, City University of Hong Kong; Hong Kong Institute of AI for Science, City University of Hong Kong