🤖 AI Summary
This study addresses the limitation of Reinforcement Learning with Verifiable Rewards (RLVR), where global entropy conflates token importance and leads to inaccurate local credit assignment. To overcome this, we propose PEPO, an algorithm that introduces a difficulty- and position-invariant proximal entropy metric to quantify local relative proximity. This metric is employed to weight the token advantage function, enabling precise credit assignment within a policy optimization framework built upon GRPO. Our contributions demonstrate that PEPO effectively mitigates the confounding effects of global entropy. Extensive experiments on mathematical reasoning tasks across multiple large language models show that PEPO significantly outperforms baseline methods while exhibiting strong generalizability.
📝 Abstract
Value-model-free RLVR methods such as GRPO assign uniform advantages to all tokens in a rollout, ignoring that tokens contribute unequally. Recent methods use token entropy as an importance proxy but compute it globally across the batch, conflating importance with prompt difficulty and positional trends. We argue that importance should instead be measured relative to the local context of each token. We introduce proximal entropy, a local measure of token importance relative to neighboring tokens, and prove it is invariant to both confounders. Proximal Entropy Policy Optimization (PEPO) uses it to weight per-token advantages and outperforms GRPO and entropy-based baselines on mathematical reasoning across Qwen3-1.7B, Qwen3-4B, and Llama-3.2-3B-Instruct. We also show the formulation generalizes to other algorithms where substituting proximal entropy into existing methods improves, and applying it to single-stream RL succeeds where global entropy fails.