🤖 AI Summary
This study addresses the training instability and inefficiency arising from critic abandonment in reinforcement learning for long chain-of-thought reasoning with large language models. We propose RFPO, a method that challenges the prevailing critic-free paradigm by repurposing a frozen pretrained critic to simultaneously serve as a reward model, baseline, and predictor. By binarizing scores to eliminate length bias and generate dense learning signals, RFPO integrates Generalized Advantage Estimation to achieve efficient policy optimization without external reward labels. Experiments demonstrate that RFPO matches the performance of supervised PPO under fully unlabeled conditions while significantly reducing computational and memory overhead. This work presents an effective reward-free post-training solution for long-horizon reasoning tasks.
📝 Abstract
Recent approaches to reinforcement learning (RL) post-training for large language models increasingly remove the critic to reduce training instability and memory overhead. Even where a critic is trained, it is discarded once training ends, although it has learned to predict outcomes. We revisit this trend and show that a pretrained critic's ability to predict future outcomes can make it a valuable asset for efficient long-horizon reasoning. First, we find that instability in critic-based RL for long chain-of-thought reasoning is largely an optimization artifact: keeping policy updates small and low in variance restores stable convergence. Second, a well-pretrained critic estimates the posterior probability of eventual success from later trajectory states and unfinished prefixes. Its predictions provide outcome-derived, dense, per-prefix learning signals that, during policy optimization, require neither completed rollouts, step-level annotations, nor external reward labels. Building on this insight, we introduce Reward-Free Policy Optimization (RFPO), which repurposes a single calibrated, frozen critic as a rollout-level reward, a value baseline for generalized advantage estimation, and a success forecaster for unfinished prefixes. We further show that binarizing the debiased score stops the policy from exploiting the critic's length bias. Binarized, RFPO matches supervised PPO without a single label in the training loop, while cutting compute and memory overhead. This makes RFPO well suited to long-horizon reasoning tasks, where outcomes arrive late and generation dominates cost: because rollouts can be rewarded before they finish, training no longer has to pay for waiting on every trajectory to complete. Our findings challenge the prevailing critic-free paradigm and establish critic-based, reward-free optimization as a scalable and computationally efficient path for LLM post-training.