Function-Structured Reinforcement Learning with Executable Verifiers for Mathematical Reasoning

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of localizing intermediate-step errors in algorithmic mathematical reasoning and the discrepancy between execution success and logical correctness by proposing the FSG-RL framework. This method introduces a novel reinforcement learning paradigm that integrates function structure graphs with executable verifiers, decomposing problems into subproblem graphs and Python code. The solving process is optimized through answer-gated rewards, span-level credit assignment, and a teacher-supervised GRPO strategy. Experimental results demonstrate that final-answer accuracy improves from 43.25% to 67.50%, while complete problem-solving success rates increase from 32.25% to 52.25%, significantly enhancing the reliability of model reasoning.
📝 Abstract
Algorithmic mathematical reasoning requires reliable decomposition, computation, and aggregation. Final-answer rewards provide limited guidance on intermediate errors, while successful execution does not guarantee mathematical correctness. This work proposes Function-Structured Graph Reinforcement Learning (FSG-RL), connecting subproblem graphs and Python implementations with multi-verifier feedback. The policy first learns to generate code from function graphs through supervised fine-tuning (SFT). Group Relative Policy Optimization (GRPO) then optimizes the policy using answer-gated rewards and span-level credit assignment. The framework also supports teacher supervision and structured memory. A benchmark curated from Grade School Math 8K (GSM8K), MathQA, MATH, and Omni-MATH pairs public function graphs with private verification specifications. Under a unified evaluation protocol, GRPO improves final-answer accuracy from 43.25% to 67.50% and full solution success from 32.25% to 52.25% over SFT. Continued reinforcement learning (RL) with teacher supervision yields additional gains. The gains extend beyond producing correctly formatted code, supporting verifier-guided reinforcement learning for mathematical reasoning. Code is available at https://github.com/ZihanLiummyycc/FSG-RL.
Problem

Research questions and friction points this paper is trying to address.

Mathematical Reasoning
Reinforcement Learning
Intermediate Error Guidance
Executable Verification
Innovation

Methods, ideas, or system contributions that make the work stand out.

Function-Structured Graph Reinforcement Learning
Group Relative Policy Optimization (GRPO)
Multi-verifier Feedback
Mathematical Reasoning
Span-level Credit Assignment