🤖 AI Summary
This work addresses the high computational cost of online post-training for large language models and the insufficient supervision provided by fixed demonstrations. We propose a rollout-free self-distillation method that departs from the conventional paradigm of treating reference solutions as undifferentiated context. Instead, it innovatively converts demonstrations into context-dependent distributional supervision signals via a step-alignment mechanism. By integrating progress-correlated reasoning transitions with privileged guidance, we construct an efficient offline self-training framework. Experiments demonstrate that our approach significantly outperforms supervised fine-tuning on mathematical reasoning benchmarks while achieving performance comparable to online reinforcement learning. Furthermore, it preserves code generation capabilities and yields approximately a 2× training speedup.
📝 Abstract
On-policy post-training can improve large language models by learning from their own trajectories, but requires costly rollout generation. We ask whether fixed demonstrations can support competitive off-policy learning through better supervision. Our premise is that their usefulness depends not only on the training trajectories, but also on whether supervision provides informative preferences among continuations and connects this guidance to the reasoning decision being learned. We introduce Step-Aligned Privileged Distillation (SAPD), a rollout-free self-distillation method that turns demonstrations into step-aligned distributional supervision. Its key insight is to use the known progression of a reference solution to associate each reasoning transition with targeted privileged guidance, rather than treating the solution as undifferentiated context. On mathematical reasoning benchmarks, SAPD outperforms supervised fine-tuning and label smoothing on average while remaining competitive with on-policy reinforcement learning and self-distillation. Analyses support both the value of context-dependent distributional guidance and the benefit of aligning privileged information with the current step. SAPD also largely preserves out-of-domain coding performance and achieves approximately 2x training-loop speedups over the on-policy baselines. These findings suggest that carefully constructed supervision can make fully off-policy post-training a competitive and computationally efficient alternative. Our code is available at https://github.com/Miaow-Lab/SAPD.