Inducing Process Supervision from Outcome-Only Reinforcement Learning

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive annotation costs and computationally expensive Monte Carlo estimation bottlenecks inherent in process reward models (PRMs). We propose TIPS, a framework that integrates a generative PRM architecture with chain-of-thought reasoning. By leveraging only outcome labels and optimizing via Group Relative Policy Optimization, TIPS elicits step-level verification capabilities without requiring explicit process supervision. Theoretically, we demonstrate that pure outcome feedback suffices to reinforce the correctness of intermediate reasoning steps. Empirically, TIPS-Qwen3-4B surpasses strong baselines such as GPT-5.4 on mathematical and agent benchmarks using minimal data, achieving an F1 score of 85.2 on ProcessBench while substantially reducing annotation overhead.
📝 Abstract
Process reward models (PRMs) have become a key component for LLMs, as their step-level feedback supports both post-training and test-time reasoning. However, training strong PRMs remains costly: human step annotation is difficult to scale, while Monte Carlo estimation is computationally expensive and can drift from the intrinsic correctness of steps. To get effective PRMs at low cost, we introduce TIPS (Thinking-Induced Process Supervision), an outcome-only reinforcement learning (RL) framework for training generative PRMs. In TIPS, the model generates a chain-of-thought (CoT) followed by step-level labels and an outcome label. The reward depends solely on whether the predicted outcome matches the ground truth, and the resulting group-relative advantage is used to optimize the entire generated response. Intuitively, when checking intermediate steps helps determine the outcome, more accurate checks can lead to better outcome judgments and higher rewards. Outcome-only RL can therefore reinforce step-level verification without explicit process supervision. We validate the effectiveness of TIPS across math and agent benchmarks and four backbone families. Notably, TIPS-Qwen3-4B-Thinking-2507 reaches 85.2 F1 on ProcessBench with only 3.2K outcome-labeled trajectories, surpassing all evaluated trained PRMs and strong prompt-only judges such as GPT-5.4-Instruct and Claude-4.7-Opus, while still trailing o1-mini. Code and data are available at https://github.com/RUCBM/TIPS.
Problem

Research questions and friction points this paper is trying to address.

Process Reward Models
Outcome-only Reinforcement Learning
Step-level Feedback
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Process Reward Models
Outcome-only Reinforcement Learning
Chain-of-Thought
Step-level Verification
Generative PRMs
🔎 Similar Papers
No similar papers found.