ReSPO: Reshaped Sequence Policy Optimization for Gradient Starvation in Off-Policy Learning

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the gradient starvation problem in off-policy learning, where importance weight clipping suppresses low-weight positive samples while allowing high-weight negative samples to dominate optimization. To mitigate this, we propose ReSPO, which replaces hard clipping with a smooth, dual-branch sequence-level kernel to balance gradient contributions from positive and negative samples. Specifically, an asymmetric kernel function is designed based on an α-divergence variational objective and exponential variance control, effectively preserving learning signals from long positive reasoning trajectories. Experiments on Qwen3 models with both Dense and Mixture-of-Experts (MoE) architectures demonstrate that ReSPO significantly accelerates early-stage convergence and improves final benchmark performance, thereby validating the effectiveness of tail weight control in off-policy optimization.
📝 Abstract
Reinforcement learning from verifiable rewards (RLVR) frequently reuses rollouts across multiple policy updates, increasing the mismatch between the current policy and the data-generating policy. We identify a sign-dependent gradient starvation problem in clipped policy optimization: clipping suppresses under-generated positive responses at the low-importance-weight tail while permitting severely over-generated negative responses to dominate the high-weight tail. To address this, we propose ReSPO (Reshaped Sequence Policy Optimization), which replaces clipping with a smooth, two-branch sequence-level kernel derived from an $\alpha$-divergence variational objective and an exponential variance-control tilt. The positive branch preserves a nonzero gradient weight for under-generated positive responses, while the negative branch suppresses heavily over-generated negative responses. We demonstrate that ReSPO effectively learns from long positive reasoning trajectories during early training, even when accumulated policy drift relegates them to the low-importance-weight tail. On dense and MoE Qwen3 models, ReSPO accelerates early optimization, improves final training scores, and achieves higher held-out benchmark performance under a rollout reuse, validating our approach on importance-weight tail control in off-policy learning.
Problem

Research questions and friction points this paper is trying to address.

Off-Policy Learning
Gradient Starvation
Reinforcement Learning from Verifiable Rewards
Policy Optimization
Importance Weight
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reshaped Sequence Policy Optimization
Gradient Starvation
Off-Policy Learning
Alpha-Divergence
Importance Weight Tail Control
Y
Yihang Chen
Department of Computer Science, University of California, Los Angeles
Y
Yuanhao Ban
Department of Computer Science, University of California, Los Angeles
Cho-Jui Hsieh
Cho-Jui Hsieh
University of California, Los Angeles
Machine LearningOptimization