Solving Without Stopping: On-Policy Distillation at Small Scale

πŸ“… 2026-09-29
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the significant imbalance in transferring reasoning and stopping judgment capabilities during knowledge distillation to smaller models. Leveraging the multi-scale Qwen3 model series, we conduct online policy distillation experiments under both thinking and non-thinking modes. We propose a diagnostic framework that disentangles answer tokens, correctness, and stopping behaviors, revealing that distillation merely preserves pre-existing correct stopping points rather than acquiring new ones. Our findings demonstrate that smaller models struggle to learn novel stopping signals and frequently over-generate due to an inability to confirm their answers. This work identifies a dual ceiling on both reasoning performance and termination decision-making in this context, offering a new perspective for understanding the scaling effects of capability transfer in knowledge distillation.
πŸ“ Abstract
On-policy distillation, where a student learns from a stronger teacher's feedback on its own outputs, is a common way to pass reasoning to smaller models. We analyze what it transfers at small scale, distilling Qwen3-8B into Qwen3 4B, 1.7B and 0.6B students, in thinking mode (reason at length, then end the reasoning and answer) and, for comparison, in non-thinking mode (no separate reasoning phase). Long reasoning needs two abilities, solving a problem and knowing when it is solved, and we find that distillation transfers the first, but in thinking mode not the second. Solving improves at every size, up to two ceilings, which we measure comprehensively across both modes and all student sizes: a student's single attempt never exceeds what it could already reach in many attempts before training, and the smaller the student, the further it stays below the teacher. Stopping is where the modes part. In non-thinking mode every student keeps stopping; in thinking mode students stop ending their reasoning early in training, and the smaller the student, the less of this ability survives: the teacher signals a stop almost only where a student already ends its reasoning, so distillation teaches no new stops; it only keeps the student's existing stops that land on a right answer, and a weak student has few such stops. The smallest students often reach the right value but do not commit to it: they either rarely mark it or mark it and write past it. Together, these results describe how small students behave under on-policy distillation, and a diagnostic that separates answer marking, correctness and stopping.
Problem

Research questions and friction points this paper is trying to address.

on-policy distillation
reasoning transfer
small language models
stopping ability
thinking mode
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-policy Distillation
Reasoning Transfer
Stopping Behavior
Thinking Mode
Diagnostic Framework
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
H
Hongyang Li
University of Luxembourg
Yiming Zhu
Yiming Zhu
Phd of AI
Social computingInternet MeasurementsData Science
X
Xiao Li
Seafill Open-Source Community
C
Caesar Wu
University of Luxembourg
S
Said Mammar
UniversitΓ© Paris-Saclay
Pascal Bouvry
Pascal Bouvry
Professor of Computer Science, University of Luxembourg
OptimisationCloud/Distributed/Parallel ComputingAd Hoc networks