On the Impossibility of Unbiased and Length-Invariant Policy Optimization with Outcome Rewards

๐Ÿ“… 2026-07-25
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work establishes an impossibility theorem in outcome-reward-based policy optimization, demonstrating that unbiasedness and trajectory-length invariance cannot be simultaneously achieved under standard assumptions. Specifically, it proves that no length-weighting scheme can satisfy both properties and fully characterizes the resulting trade-off spectrum. By introducing a parameterized family of weighting functions \( f_\alpha(L) = L^{\alpha - 1} \), the framework unifies GRPO (\(\alpha=0\)) and Dr. GRPO (\(\alpha=1\)): the former preserves length invariance but yields biased gradients, while the latter ensures unbiasedness at the cost of heightened sensitivity to long trajectories, with gradient contributions scaling linearly with trajectory length. The analysis reveals that these two algorithms occupy opposite ends of the biasโ€“invariance trade-off, implying no universally optimal solution exists.
๐Ÿ“ Abstract
Group Relative Policy Optimization (GRPO) is the dominant reinforcement learning algorithm for training reasoning capabilities in large language models, notably adopted by DeepSeek-R1. The recent improvement Dr. GRPO (COLM 2025) identifies the response-level length bias caused by per-trajectory length normalization in GRPO and proposes removing this normalization, claiming the resulting optimizer is "unbiased." We show that this claim is incomplete. Specifically, we establish an impossibility theorem: under the standard outcome reward + GRPO setting, no length-based weighting scheme can simultaneously achieve the following two properties. (P1) Gradient unbiasedness: the gradient estimator is an unbiased estimate of the true policy gradient. (P2) Length invariance: each trajectory's effective contribution to the gradient is independent of its token length. GRPO approximately satisfies P2 but violates P1; Dr. GRPO satisfies P1 but violates P2. We characterize the complete tradeoff spectrum via the parametric family f_alpha(L) = L^{alpha - 1}, where alpha = 0 recovers GRPO, alpha = 1 recovers Dr. GRPO, and provide quantitative analysis showing that Dr. GRPO's length bias can cause longer trajectories to dominate gradient updates by a factor proportional to the length ratio. Our results reveal that neither algorithm is universally "done right"; they occupy opposite ends of a fundamental and unavoidable tradeoff.
Problem

Research questions and friction points this paper is trying to address.

policy optimization
outcome rewards
length bias
gradient unbiasedness
length invariance
Innovation

Methods, ideas, or system contributions that make the work stand out.

impossibility theorem
policy optimization
length bias
gradient unbiasedness
outcome rewards
๐Ÿ”Ž Similar Papers
Fei Ding
Fei Ding
Unknown affiliation
Y
Yongkang Zhang
Alibaba Group
Y
Yuhao Liao
Tsinghua University
Z
Zijian Zeng
Tsinghua University
H
Huiming Yang
Tsinghua University