🤖 AI Summary
This study addresses the problem of inaccurate step-level credit assignment in long-horizon agent reinforcement learning by proposing a cross-trajectory Bellman closure method. This approach propagates evidence through the linear solution of empirical process fixed points, effectively integrating the advantages of local averaging and global propagation. By incorporating group-based reinforcement learning, behavioral policy Bellman fixed-point solving, and finite-depth estimation techniques, the method generates precise step-level credit signals for policy optimization without requiring additional environment interactions. Experimental results demonstrate that the proposed method achieves substantial performance improvements on benchmarks such as ALFWorld, outperforming the strongest baseline by 5.59 percentage points.
📝 Abstract
Group-based reinforcement learning such as GRPO trains LLM agents by comparing rollouts sampled for each task, without a learned critic. In long-horizon settings, these rollouts revisit shared anchor states, offering cross-rollout evidence for step-level credit. Ideally, step-level credit should incorporate evidence beyond the realized suffixes observed at an anchor while aggregating alternative continuations according to their empirical frequencies. Visit-local averaging pools realized suffix returns at shared anchors and respects observed frequencies, but does not recursively propagate evidence across rollouts, whereas shortest-path estimators have global reach but allow a rarely observed route to dominate an anchor's value. We introduce Cross-Rollout Bellman Closure (CRBC), which merges each rollout group into a finite empirical process with absorbing success and failure boundaries and evaluates its behavior-policy Bellman fixed point with one linear solve. This fixed point uses the same empirical action and transition frequencies to propagate evidence through shared anchors and aggregate alternative continuations. Backing up the resulting state values through observed transitions yields action values, whose gain over the corresponding state value provides step-level credit. A corresponding finite-depth family recovers visit-local return averaging at zero depth and converges to the exact closure as depth increases. The normalized closure credit is combined with the trajectory-level group advantage for policy optimization, without additional environment rollouts. Across ALFWorld, WebShop, and Sokoban benchmarks with multiple model scales, CRBC consistently improves final performance and learning efficiency. For example, CRBC outperforms the strongest evaluated baseline by 5.59 percentage points on ALFWorld with Qwen2.5-1.5B-Instruct.