🤖 AI Summary
This study addresses the limitation in reinforcement learning where high-level policies are evaluated solely based on final outcomes, resulting in insufficient signal for policy quality. To overcome this, we propose the SURE framework, which decouples strategy from execution by formally defining and learning "policy utility" for the first time. Specifically, a reward model is trained using preference data generated via teacher-model comparisons and subsequently integrated into the GRPO algorithm to enable policy-level optimization. Evaluated across three base models, our approach yields average improvements of 1.87%–2.93% in pass@1 accuracy on mathematical reasoning benchmarks while significantly reducing computational overhead. This work establishes a new paradigm for efficient policy evaluation under constrained computational resources.
📝 Abstract
Reinforcement learning with verifiable rewards has substantially improved mathematical reasoning. However, terminal correctness alone provides limited insight into the quality of high-level strategies, such as theorem selection and subgoal decomposition, when considered separately from their subsequent execution. This paper studies strategy utility, which is defined as the likelihood that a strategy supports a correct downstream solution under a given executor. We introduce SURE, a framework for learning and leveraging relative strategy utility. In this framework, high-level strategies are separated from their detailed reasoning. Based on the pairwise preferences constructed from strategy-conditioned rollouts and teacher-generated contrasts, a Strategy Reward Model is learned to estimate relative strategy utility. During reinforcement learning, the frozen reward model reads only the extracted strategy, whose score is combined with the correctness and format rewards in a sequence-level GRPO objective. Compared with outcome-and-format GRPO baselines, experiments show that SURE improves average pass@1 by 1.87%, 2.64%, and 2.93% across three policy backbones. Our method also achieves competitive or better accuracy than stronger reward baselines while requiring substantially lower GRPO-stage compute.