🤖 AI Summary
This work addresses the limited reasoning capabilities of large language models (LLMs) by deeply integrating reinforcement learning (RL) into the entire “slow-thinking” stepwise reasoning process—both during training and decoding. We introduce the first formalized “slow-thinking” reasoning framework, unifying model-based (e.g., Monte Carlo Tree Search) and model-free (e.g., Proximal Policy Optimization) RL paradigms to render reasoning steps learnable, intervenable, and interpretable. Our method jointly incorporates stepwise reasoning modeling, chain-of-thought fine-tuning, reward modeling over reasoning trajectories, and in-model policy optimization. Evaluated on benchmarks including GSM8K and MMLU, our approach achieves 12–18% absolute improvements in mathematical reasoning and complex problem-solving over strong baselines. Moreover, generated reasoning processes exhibit enhanced reliability, transparency, and traceability—enabling rigorous auditing and human-in-the-loop refinement.
📝 Abstract
OpenAI o1 has shown that applying reinforcement learning to integrate reasoning steps directly during inference can significantly improve a model's reasoning capabilities. This result is exciting as the field transitions from the conventional autoregressive method of generating answers to a more deliberate approach that models the slow-thinking process through step-by-step reasoning training. Reinforcement learning plays a key role in both the model's training and decoding processes. In this article, we present a comprehensive formulation of reasoning problems and investigate the use of both model-based and model-free approaches to better support this slow-thinking framework.