๐ค AI Summary
This work establishes an impossibility theorem in outcome-reward-based policy optimization, demonstrating that unbiasedness and trajectory-length invariance cannot be simultaneously achieved under standard assumptions. Specifically, it proves that no length-weighting scheme can satisfy both properties and fully characterizes the resulting trade-off spectrum. By introducing a parameterized family of weighting functions \( f_\alpha(L) = L^{\alpha - 1} \), the framework unifies GRPO (\(\alpha=0\)) and Dr. GRPO (\(\alpha=1\)): the former preserves length invariance but yields biased gradients, while the latter ensures unbiasedness at the cost of heightened sensitivity to long trajectories, with gradient contributions scaling linearly with trajectory length. The analysis reveals that these two algorithms occupy opposite ends of the biasโinvariance trade-off, implying no universally optimal solution exists.
๐ Abstract
Group Relative Policy Optimization (GRPO) is the dominant reinforcement learning algorithm for training reasoning capabilities in large language models, notably adopted by DeepSeek-R1. The recent improvement Dr. GRPO (COLM 2025) identifies the response-level length bias caused by per-trajectory length normalization in GRPO and proposes removing this normalization, claiming the resulting optimizer is "unbiased." We show that this claim is incomplete. Specifically, we establish an impossibility theorem: under the standard outcome reward + GRPO setting, no length-based weighting scheme can simultaneously achieve the following two properties. (P1) Gradient unbiasedness: the gradient estimator is an unbiased estimate of the true policy gradient. (P2) Length invariance: each trajectory's effective contribution to the gradient is independent of its token length. GRPO approximately satisfies P2 but violates P1; Dr. GRPO satisfies P1 but violates P2. We characterize the complete tradeoff spectrum via the parametric family f_alpha(L) = L^{alpha - 1}, where alpha = 0 recovers GRPO, alpha = 1 recovers Dr. GRPO, and provide quantitative analysis showing that Dr. GRPO's length bias can cause longer trajectories to dominate gradient updates by a factor proportional to the length ratio. Our results reveal that neither algorithm is universally "done right"; they occupy opposite ends of a fundamental and unavoidable tradeoff.