🤖 AI Summary
This study addresses the training collapse problem in Proximal Policy Optimization (PPO) for large language model reinforcement learning, which is primarily caused by critic instability. To mitigate this issue, we propose a comprehensive set of critic stabilization mechanisms. Methodologically, we stabilize critic updates through excessive-length filtering, noise-normalized regression, and reduced batch sizes applied to the actor. Furthermore, we introduce truncated bias correction and heterogeneous noise handling strategies to optimize gradient clipping, while enhancing the PPO algorithm via statistical weighting and dynamic batching techniques. Experimental results demonstrate that the proposed approach achieves consistently stable training throughout the entire process on coding and mathematical reasoning tasks, improving verification scores by up to 14.89% compared to standard PPO.
📝 Abstract
A key strength of Proximal Policy Optimization (PPO) is its learned critic, which uses historical trajectories collected during reinforcement learning to estimate expected returns and reduce policy-gradient variance. However, we find that the critic is also a major source of instability in reinforcement learning for large language models (LLMs). We identify two critic failure modes that destabilize PPO. First, filtering truncated rollouts from both actor and critic shifts the policy objective to reward conditioned on completion, allowing truncation to increase even as conditional reward improves. Second, heterogeneous return noise can cause high-variance prompts to dominate critic updates in finite batches. We introduce EasyPPO to address these failures. Actor-only overlong filtering trains the critic on returns from both completed and truncated rollouts. Noise-normalized critic regression weights each prompt's critic loss by the inverse standard deviation of its sampled returns, balancing noise contributions across prompts. Moderately smaller critic mini-batches confine outlier influence to fewer rollouts during gradient clipping. Across continuous-reward coding on FrontierCS, binary-reward mathematical reasoning on AIME24, and multi-turn search on Search-R1, EasyPPO remains stable throughout the full training horizon and consistently outperforms vanilla PPO, VAPO, and HL-Gauss PPO. Its best validation scores show relative gains of 14.89%, 2.28%, and 9.47% over PPO, respectively.