🤖 AI Summary
This study addresses the gradient starvation problem in off-policy learning, where importance weight clipping suppresses low-weight positive samples while allowing high-weight negative samples to dominate optimization. To mitigate this, we propose ReSPO, which replaces hard clipping with a smooth, dual-branch sequence-level kernel to balance gradient contributions from positive and negative samples. Specifically, an asymmetric kernel function is designed based on an α-divergence variational objective and exponential variance control, effectively preserving learning signals from long positive reasoning trajectories. Experiments on Qwen3 models with both Dense and Mixture-of-Experts (MoE) architectures demonstrate that ReSPO significantly accelerates early-stage convergence and improves final benchmark performance, thereby validating the effectiveness of tail weight control in off-policy optimization.
📝 Abstract
Reinforcement learning from verifiable rewards (RLVR) frequently reuses rollouts across multiple policy updates, increasing the mismatch between the current policy and the data-generating policy. We identify a sign-dependent gradient starvation problem in clipped policy optimization: clipping suppresses under-generated positive responses at the low-importance-weight tail while permitting severely over-generated negative responses to dominate the high-weight tail. To address this, we propose ReSPO (Reshaped Sequence Policy Optimization), which replaces clipping with a smooth, two-branch sequence-level kernel derived from an $\alpha$-divergence variational objective and an exponential variance-control tilt. The positive branch preserves a nonzero gradient weight for under-generated positive responses, while the negative branch suppresses heavily over-generated negative responses. We demonstrate that ReSPO effectively learns from long positive reasoning trajectories during early training, even when accumulated policy drift relegates them to the low-importance-weight tail. On dense and MoE Qwen3 models, ReSPO accelerates early optimization, improves final training scores, and achieves higher held-out benchmark performance under a rollout reuse, validating our approach on importance-weight tail control in off-policy learning.