Dual-Channel Robust Group-Relative Policy Optimization via Advantage and Sequence-Weight Estimation

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the susceptibility of reward and sequence weights to outlier interference in Group Relative Policy Optimization (GRPO), which leads to contrast collapse and clipping bias. To mitigate these issues, this work proposes RoVR-GSPO, a dual-channel robust optimization framework. The method introduces a novel architecture that decouples the reward and ratio channels, integrating robust reference estimation, bounded residual credit assignment, and differentiable SoftRoVR aggregation to independently suppress anomalies across both channels. Experimental results demonstrate that the proposed framework significantly outperforms GSPO on tasks such as mathematical reasoning while exhibiting strong resilience under data perturbations. Ultimately, RoVR-GSPO effectively enhances the stability of policy optimization, offering a principled solution for reliable reinforcement learning from human feedback.
📝 Abstract
Group-relative policy optimization relies on reward-derived advantages and sequence-level likelihood weights, both of which can be sensitive to localized outliers. Extreme rewards can collapse the contrast among clean responses after group normalization, while token-level log-ratio perturbations can alter sequence weights and clipping decisions. We introduce RoVR-GSPO, a dual-channel robust optimizer that addresses these failure modes separately. Its reward channel combines robust reference estimation with bounded residual credit, while its ratio channel uses differentiable SoftRoVR aggregation to construct robust sequence weights. We provide stability and efficiency analyses for both channels. Experiments on mathematical reasoning, long-context summarization, and tool-call annotation show consistent improvements over GSPO, while controlled perturbation studies demonstrate stronger robustness to reward contamination and token-ratio anomalies.
Problem

Research questions and friction points this paper is trying to address.

Group-Relative Policy Optimization
Reward Outliers
Sequence Weights
Robustness
Token-level Perturbations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dual-Channel Optimization
Robust Policy Optimization
Group-Relative Policy Optimization
SoftRoVR Aggregation
Bounded Residual Credit
🔎 Similar Papers