From Reasoning Chains to Verifiable Subproblems: Curriculum Reinforcement Learning Enables Credit Assignment for LLM Reasoning

📅 2026-05-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge that large language models struggle to learn from sparse correct answers in complex reasoning tasks, as conventional outcome-based rewards fail to leverage partial progress from failed attempts. To overcome this, the authors propose the SCRL framework, which constructs a curriculum of verifiable subproblems derived from reference reasoning chains, treats the original problem as the final subproblem, and applies reward normalization at the subproblem level to enable fine-grained credit assignment. SCRL is the first approach to integrate verifiable subproblem curricula with subproblem-level reward normalization, transforming partial reasoning progress into effective learning signals without external scoring and mitigating gradient vanishing in hard problems. Evaluated on seven mathematical reasoning benchmarks, SCRL significantly outperforms strong baselines, achieving an average accuracy gain of 4.1 points on Qwen3-4B-Base and improvements of 3.7 and 4.6 points in pass@1 and pass@64 on AIME24/25 and IMO-Bench, respectively.
📝 Abstract
Reinforcement learning from verifiable rewards (RLVR) has shown strong promise for LLM reasoning, but outcome-based RLVR remains inefficient on hard problems because correct final-answer rollouts are rare and sample-level credit assignment cannot use partial progress in failed attempts. We introduce SCRL (Subproblem Curriculum Reinforcement Learning), a curriculum RL framework that derives verifiable subproblems from reference reasoning chains and fixes the final subproblem as the original problem. This turns partial progress on hard problems into verifiable learning signals. Algorithmically, SCRL uses subproblem-level normalization, which normalizes rewards independently at each subproblem position and assigns the resulting advantages to the corresponding answer spans, enabling finer-grained credit assignment without external rubrics or reward models. Our analysis shows that subproblem curricula lift hard problems out of gradient dead zones, with larger relative gains as the original problem becomes harder. Across seven mathematical reasoning benchmarks, SCRL outperforms strong curriculum-learning baselines, improving average accuracy over GRPO by +4.1 points on Qwen3-4B-Base and +1.9 points on Qwen3-14B-Base. On AIME24, AIME25, and IMO-Bench, SCRL further improves pass@1 by +3.7 points and pass@64 by +4.6 points on Qwen3-4B-Base, indicating better exploration on hard reasoning problems.
Problem

Research questions and friction points this paper is trying to address.

credit assignment
LLM reasoning
reinforcement learning
verifiable subproblems
hard reasoning problems
Innovation

Methods, ideas, or system contributions that make the work stand out.

Subproblem Curriculum
Reinforcement Learning
Credit Assignment
Verifiable Rewards
LLM Reasoning
X
Xitai Jiang
LeapLab, Tsinghua University; Qiuzhen College, Tsinghua University
Z
Zihan Tang
Qiuzhen College, Tsinghua University
W
Wenze Lin
LeapLab, Tsinghua University; Qiuzhen College, Tsinghua University
Y
Yang Yue
LeapLab, Tsinghua University
S
Shenzhi Wang
LeapLab, Tsinghua University
G
Gao Huang
LeapLab, Tsinghua University