π€ AI Summary
This study addresses the performance degradation of online self-distillation in large language models as scale increases, along with the imitation gap caused by unverified trajectories. To overcome these challenges, this work proposes OASIS, a framework that decouples scaffolding from contextual roles. Unlike conventional Online Process Supervised Distillation (OPSD), which relies on complete solution steps, OASIS supervises verified online trajectories using only final answer labels and substitutes reference solutions with the modelβs own generation attempts as teacher contexts. Experimental results demonstrate that OASIS yields average improvements of 3.2 to 3.8 points across the Qwen3 model series, achieving a 3.05-point gain over OPSD for the 8B model. These findings indicate that the proposed approach effectively mitigates scaling bottlenecks while substantially enhancing reasoning accuracy.
π Abstract
On-policy self-distillation (OPSD) trains a student to match a privileged teacher distribution along its own sampled trajectory. Standard OPSD applies this supervision to unverified student rollouts while conditioning the teacher on privileged context, typically a reference solution. We separate these roles in a factorial analysis and find that scaffold correctness has a stronger effect on downstream accuracy than context correctness. Unverified scaffolds create an imitation gap because the teacher can use information unavailable to the student. This gap shrinks with model scale, yet OPSD continues to supervise mostly unverified trajectories. In contrast, verified scaffolds remain effective even when the teacher is conditioned on the student's own unsuccessful rollout. Based on this finding, we introduce OASIS, which retains the OPSD objective but supervises mostly verified by label on-policy trajectories and replaces written solutions with unverified model-generated attempts as the teacher context. OASIS therefore requires only final-answer labels. Across Qwen3-1.7B, 4B, and 8B on AIME 2024, AIME 2025, and HMMT 2025, OASIS improves over the base model by 3.2--3.8 points on average, while OPSD's gain falls from 3.05 points at 1.7B to 0.14 at 8B. At 8B, OASIS improves over OPSD by 3.05 points, showing that verified on-policy scaffolds preserve the effectiveness of self-distillation as models scale.