GRPO and Reflection Reward for Mathematical Reasoning in Large Language Models

πŸ“… 2026-03-14
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the limitation of current large language models in mathematical reasoning due to the absence of an effective active reflection mechanism, which hinders their self-correction and reasoning capabilities. To overcome this, we propose a four-stage training framework that explicitly incorporates reflection rewards during training for the first time. Our approach synergistically optimizes cognitive and environmental interaction rewards through Group Relative Policy Optimization (GRPO), while integrating both accuracy and format-based reward signals. Full-parameter supervised fine-tuning is employed to enhance the model’s introspective capacity. Experimental results demonstrate state-of-the-art performance on mathematical reasoning benchmarks. Ablation studies confirm the critical role of reflection rewards and further show that full-parameter fine-tuning significantly outperforms parameter-efficient alternatives such as LoRA.

Technology Category

Natural Language Processing: (Large) Language ModelsMachine Learning: Large Multimodal Models (LMMs)Search and Optimization: Learning to Search

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
πŸ“ Abstract
The enhancement of reasoning capabilities in large language models (LLMs) has garnered significant attention, with supervised fine-tuning (SFT) and reinforcement learning emerging as dominant paradigms. While recent studies recognize the importance of reflection in reasoning processes, existing methodologies seldom address proactive reflection encouragement during training. This study focuses on mathematical reasoning by proposing a four-stage framework integrating Group Relative Policy Optimization (GRPO) with reflection reward mechanisms to strengthen LLMs' self-reflective capabilities. Besides, this approach incorporates established accuracy and format reward. Experimental results demonstrate GRPO's state-of-the-art performance through reflection-encouraged training, with ablation studies confirming the reflection reward's pivotal role. Comparative evaluations demonstrate full-parameter SFT's superiority over low-rank adaptation (LoRA) despite heightened computational demands. Building on these cumulative findings, this research substantiates GRPO's methodological significance in post-training optimization and envisions its potential to serve as a pivotal enabler for future LLM-based intelligent agents through the synergistic integration of cognitive rewards with dynamic environmental interactions.
Problem

Research questions and friction points this paper is trying to address.

mathematical reasoning
reflection
large language models
reinforcement learning
self-reflection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Group Relative Policy Optimization
reflection reward
mathematical reasoning
self-reflection
reinforcement learning
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.