🤖 AI Summary
This study addresses the scalability limitations of on-policy self-distillation arising from its reliance on external privileged information, as well as its insufficient self-improvement when such information is unavailable. To overcome these challenges, this work proposes a self-evolution framework that eliminates the need for external annotations. By leveraging the multiple reasoning modes of large language models, the method generates Chain-of-Thought (CoT) rationales through deep thinking to serve as internal supervision signals, which subsequently guide non-thinking outputs. Furthermore, token-level supervision is employed to facilitate cross-mode knowledge distillation and shared parameter updates. Experimental results demonstrate that this framework significantly enhances dual-mode reasoning capabilities across various models and tasks, effectively surpassing the performance bottlenecks inherent in conventional self-improvement approaches.
📝 Abstract
Recent advances in online policy self-distillation (OPSD) have demonstrated that large language models (LLMs) can improve their capabilities by leveraging external privileged information (PI), such as manual annotations or feedback from external environments. However, obtaining accurate annotations and constructing sophisticated environments often require substantial human effort and computation, limiting the scalability of OPSD. While a few recent studies have explored self-improvement without external PI, the resulting gains remain limited. In this work, we explore whether LLMs can achieve comparable self-improvement without external PI. Our key observation is that a single LLM can support multiple reasoning modes, such as deep-thinking and non-thinking modes, with deep thinking generating additional information during reasoning. Based on this observation, we propose Self-Evolving Online Policy Distillation (SeOPD), which enables LLMs to distill and internalize information generated by their own chain of thought (CoT). Specifically, it (1) generates CoT with the deep-thinking mode, (2) produces responses with the non-thinking mode, and (3) uses the generated CoT as PI to provide token-level supervision for the non-thinking response, allowing new information inferred during reasoning to guide the non-thinking mode and be internalized into the shared model parameters, thereby improving both non-thinking and deep-thinking capabilities. Extensive experiments across LLMs and tasks demonstrate the effectiveness of SeOPD.