🤖 AI Summary
This study addresses the failure of supervision in token-level distillation for long-horizon reasoning, where correcting incoherence and degenerate prefixes proves ineffective. To overcome this limitation, we propose Seg-OPD, a segment-level online process distillation method. Seg-OPD dynamically selects student reasoning segments based on uncertainty measures and leverages teacher-model rewrites as explicit corrective supervision signals. By applying controlled interventions to optimize intermediate reasoning steps, it enhances the accuracy of subsequent generation. Experimental results demonstrate that Seg-OPD achieves an average relative improvement of 5.22% in reasoning accuracy across mathematical and programming tasks, significantly outperforming existing state-of-the-art baselines. The source code has been made publicly available.
📝 Abstract
On-policy distillation (OPD) improves large language model reasoning by training students on their own rollouts with dense token-wise supervision from the teacher. However, token-wise OPD does not explicitly provide a coherent alternative reasoning step showing how the student's step could be revised to improve subsequent reasoning. Furthermore, this paradigm can become less effective when the student produces a degenerate reasoning prefix, as subsequent teacher supervision remains conditioned on that prefix and may reinforce poor reasoning patterns. In this work, we focus on learning reasoning revision with segment-wise OPD to rework intermediate reasoning steps and better support subsequent reasoning. Through controlled reasoning interventions, we find that replacing student segments with teacher redrafts improves subsequent reasoning accuracy. Therefore, we address the problem of turning teacher redrafts into explicit supervision for learning to revise reasoning. We propose Segment-wise On-Policy Distillation (Seg-OPD), which selects student segments based on an uncertainty metric and obtains corresponding teacher redrafts. Seg-OPD trains the student to prefer teacher redrafts over their paired student segments while retaining dense token-wise OPD supervision. Extensive experiments on mathematical reasoning and competitive programming tasks show that Seg-OPD-trained students achieve higher revision success rates than baselines. Seg-OPD consistently outperforms the compared state-of-the-art baselines in reasoning accuracy with an average relative improvement of 5.22% across diverse models and tasks. Code is available at https://anonymous.4open.science/r/Seg-OPD.