🤖 AI Summary
This study addresses the challenge in Group Relative Policy Optimization (GRPO) for agent reinforcement learning, where distinguishing critical from minor decisions leads to ambiguous credit assignment. To overcome this, we propose ProVer, a framework that employs an LLM-based judge to identify potentially critical decision segments and validates their advantages through terminal success rate disparities, thereby achieving fine-grained credit assignment. By integrating model-based judgment with outcome-oriented advantage estimation, this design circumvents exhaustive evaluation of intermediate states, substantially improving computational efficiency. Experimental results demonstrate that ProVer achieves state-of-the-art performance on benchmarks such as ALFWorld. Notably, it yields 7%–10% improvements over GRPO across the Qwen3.5 model series while introducing only marginal generation overhead.
📝 Abstract
Group Relative Policy Optimization (GRPO) has become a promising approach for training large language model agents. However, its uniform assignment of trajectory-level advantages to all policy tokens fails to distinguish consequential decisions from less relevant ones, obscuring which intermediate decisions contributed to success. We introduce ProVer, a framework that targets potentially pivotal decisions for fine-grained credit assignment in agentic reinforcement learning. Given a rollout group, an agentic judge contrasts successful and failed trajectories to propose a segment potentially responsible for their divergent outcomes. Rather than directly trusting the judge's assessment, ProVer verifies the proposed segment by estimating its advantage from the difference in terminal success rates between current-policy continuations sampled before and after the segment. Positive estimates are then incorporated into the GRPO advantages of policy tokens within the proposed segment. By using model judgment only to select where to verify, ProVer grounds local credit in observed outcomes without exhaustively evaluating every intermediate state. Across ALFWorld, WebShop, and SearchQA, ProVer achieves the strongest average performance at both model scales, with relative improvements over GRPO of 9.91% and 7.12% for Qwen3.5-2B and Qwen3.5-4B, respectively. Further analyses demonstrate that informed segment selection improves policy training with modest additional generation overhead, even without a frontier-scale judge model, highlighting the effectiveness and efficiency of selectively targeting pivotal decisions for fine-grained credit assignment in agentic reinforcement learning.