UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma

๐Ÿ“… 2026-07-08
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses a critical limitation in reinforcement learning: while importance sampling clipping ensures training stability, it inadvertently suppresses exploration of low-confidence yet correct reasoning trajectories. To overcome this, the authors propose UP (Universal Plug-in Objective), a general-purpose optimization framework featuring asymmetric gradient designโ€”unclipped gradients are preserved in positive advantage directions to enhance exploration, while clipped gradients in negative directions maintain stability. The study introduces the novel concept of probability capacity to formally characterize the structural constraints clipping imposes on exploration and develops an asymmetric, unbounded forward optimization framework using stop-gradient operators, thereby transcending conventional policy update budget limitations. UP supports both token-level (e.g., GRPO, DAPO) and sequence-level (GSPO) optimization, demonstrating consistent improvements in exploration efficacy and reasoning accuracy across diverse RL algorithms, model architectures (Dense and MoE), and modalities (language and multimodal), confirming its broad applicability.
๐Ÿ“ Abstract
Reinforcement learning (RL) has become the standard paradigm for enhancing the complex reasoning capabilities of large language models (LLMs). To achieve sample efficiency, modern RL frameworks rely on importance sampling (IS). However, these algorithms suffer from an exploration-stability dilemma. Pure IS often leads to catastrophic training instability, while standard clipping mechanisms used to mitigate this instability strictly constrain the policy update budget. By formalizing the concept of Probability Capacity (Cap), we reveal that conservative clipping structurally stifles exploration by prematurely truncating the update budget for correct but low-confidence reasoning paths. To break free from these constraints, we propose Unbounded Positive Asymmetric Optimization (UP), a universal and plug-and-play objective. UP theoretically restructures the optimization process by anchoring the policy to its current state via the stop-gradient operator. This asymmetric design unleashes unclipped, stable gradients for positive advantages to maximize exploration, while maintaining standard clipping safeguards for negative advantages to prevent training instability. Furthermore, our formulation readily extends across different optimization granularities, including token-level (GRPO, DAPO) and sequence-level (GSPO) frameworks. Extensive experiments demonstrate that UP enhances exploration capacity and achieves superior reasoning accuracy across diverse RL algorithms (DAPO, GSPO, and GRPO), model architectures (Dense, MoE, and vision-language), and training modalities (language and multimodal), validating UP as a truly universal plug-and-play enhancement for RL-based training.
Problem

Research questions and friction points this paper is trying to address.

exploration-stability dilemma
importance sampling
policy update budget
reinforcement learning
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unbounded Positive Asymmetric Optimization
Exploration-Stability Dilemma
Probability Capacity
Stop-Gradient Anchoring
Plug-and-Play RL Objective
๐Ÿ”Ž Similar Papers
2024-07-09Neural Information Processing SystemsCitations: 3