π€ AI Summary
This work addresses the limitation of current large language models in mathematical reasoning due to the absence of an effective active reflection mechanism, which hinders their self-correction and reasoning capabilities. To overcome this, we propose a four-stage training framework that explicitly incorporates reflection rewards during training for the first time. Our approach synergistically optimizes cognitive and environmental interaction rewards through Group Relative Policy Optimization (GRPO), while integrating both accuracy and format-based reward signals. Full-parameter supervised fine-tuning is employed to enhance the modelβs introspective capacity. Experimental results demonstrate state-of-the-art performance on mathematical reasoning benchmarks. Ablation studies confirm the critical role of reflection rewards and further show that full-parameter fine-tuning significantly outperforms parameter-efficient alternatives such as LoRA.
π Abstract
The enhancement of reasoning capabilities in large language models (LLMs) has garnered significant attention, with supervised fine-tuning (SFT) and reinforcement learning emerging as dominant paradigms. While recent studies recognize the importance of reflection in reasoning processes, existing methodologies seldom address proactive reflection encouragement during training. This study focuses on mathematical reasoning by proposing a four-stage framework integrating Group Relative Policy Optimization (GRPO) with reflection reward mechanisms to strengthen LLMs' self-reflective capabilities. Besides, this approach incorporates established accuracy and format reward. Experimental results demonstrate GRPO's state-of-the-art performance through reflection-encouraged training, with ablation studies confirming the reflection reward's pivotal role. Comparative evaluations demonstrate full-parameter SFT's superiority over low-rank adaptation (LoRA) despite heightened computational demands. Building on these cumulative findings, this research substantiates GRPO's methodological significance in post-training optimization and envisions its potential to serve as a pivotal enabler for future LLM-based intelligent agents through the synergistic integration of cognitive rewards with dynamic environmental interactions.