🤖 AI Summary
This study addresses the susceptibility of reward and sequence weights to outlier interference in Group Relative Policy Optimization (GRPO), which leads to contrast collapse and clipping bias. To mitigate these issues, this work proposes RoVR-GSPO, a dual-channel robust optimization framework. The method introduces a novel architecture that decouples the reward and ratio channels, integrating robust reference estimation, bounded residual credit assignment, and differentiable SoftRoVR aggregation to independently suppress anomalies across both channels. Experimental results demonstrate that the proposed framework significantly outperforms GSPO on tasks such as mathematical reasoning while exhibiting strong resilience under data perturbations. Ultimately, RoVR-GSPO effectively enhances the stability of policy optimization, offering a principled solution for reliable reinforcement learning from human feedback.
📝 Abstract
Group-relative policy optimization relies on reward-derived advantages and sequence-level likelihood weights, both of which can be sensitive to localized outliers. Extreme rewards can collapse the contrast among clean responses after group normalization, while token-level log-ratio perturbations can alter sequence weights and clipping decisions. We introduce RoVR-GSPO, a dual-channel robust optimizer that addresses these failure modes separately. Its reward channel combines robust reference estimation with bounded residual credit, while its ratio channel uses differentiable SoftRoVR aggregation to construct robust sequence weights. We provide stability and efficiency analyses for both channels. Experiments on mathematical reasoning, long-context summarization, and tool-call annotation show consistent improvements over GSPO, while controlled perturbation studies demonstrate stronger robustness to reward contamination and token-ratio anomalies.