Luck Is Not Skill: When Do Paired Rollouts Help Group-Relative RL of LLM Agents?

📅 2026-09-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过配对回放方法减少环境噪声对组相对强化学习的影响,以提高大型语言模型代理的性能。
📝 Abstract
Group-relative reinforcement learning compares rollouts of the same prompt, but independent environment noise can obscure these comparisons. We study paired rollouts, which share an event-keyed noise schedule within each group while preserving each rollout's marginal distribution. Pairing removes the between-schedule component of reward-contrast variance, but need not reduce gradient variance. For one-sided grader noise, we derive an exact condition for reduction and give a counterexample in which reward contrasts improve while gradient variance increases. A controlled study trains a 2B tool-use agent under tool faults and grader flips, with three seeds per design. The protocol was registered with a disclosed, previously completed pilot. Under tool faults, pairing improves final noisy-test success by +5.1 percentage points on average, with all three seed differences positive, but misses the registered learning-curve criterion. The criterion is also missed under grader flips: the validation-AUC difference is +0.003 (95% interval [-0.029, +0.033]). A gradient probe on eight distinct checkpoints from two fault-trained trajectories finds lower mean-centered covariance traces under both noise types: 21 to 30% for grader flips and 40 to 63% for tool faults. These finite-sample measurements support the variance mechanism without establishing a general learning-speed benefit. The results distinguish improving reward comparisons, reducing estimator variance, and improving learning.
Problem

Research questions and friction points this paper is trying to address.

group-relative reinforcement learning
paired rollouts
environment noise
reward-contrast variance
gradient variance
Innovation

Methods, ideas, or system contributions that make the work stand out.

paired rollouts
reward-contrast variance
gradient variance
group-relative reinforcement learning
🔎 Similar Papers
No similar papers found.