🤖 AI Summary
This study addresses the systematic bias in step-level advantage estimation inherent to group-based reinforcement learning methods such as GRPO when applied to agent tasks. To this end, we propose GRAFT, a framework that models trajectories as graph structures, recovers node state values via Bellman iteration, and achieves faithful credit assignment through edge differences. Furthermore, GRAFT incorporates a Graph Generalized Advantage Estimator (Graph GAE) to mitigate valuation bias and efficiently approximate multi-action sampling. Extensive evaluations on multi-turn agent benchmarks demonstrate that GRAFT consistently outperforms GRPO and other recent mainstream algorithms. By effectively resolving the challenge of fine-grained advantage estimation in complex interactive scenarios, this work provides a robust solution for improving reinforcement learning in agentic settings.
📝 Abstract
Group-based reinforcement learning (RL) methods, such as GRPO and its variants, have become a leading paradigm for training reasoning and agentic large language models (LLMs). While their group-normalized advantage estimation is reliable at the response level, it becomes systematically biased at the step level, since coarse-grained trajectory-level advantages are hard to accurately reflect the contribution of individual steps (i.e, failed trajectories may contain valuable steps). Revisiting the foundational RL definition, we notice that GRPO's success on single-turn tasks stems from its advantage estimation strategy, which adheres to the basic definition: the mean reward of multiple actions sampled from the same state constitutes a credible state-value estimate. Extending the faithful estimation to step-level would in principle demand sampling multiple actions from each intermediate state, which is too costly on a per-state basis. To mitigate this issue, we propose a Graph-based Faithful sTep-level credit-assignment framework (GRAFT) that grafts all rollout trajectories into a trajectory graph, recovering node state-values via Bellman iteration on the graph, and assigning credit to each edge by the node value difference. Theoretically, the estimated step-level advantage faithfully adheres to the basic advantage definition in RL. To further ensure the reliability of step-level advantage estimation, we further propose Graph GAE, which extends GAE to the trajectory graph for reducing the impact of state-value estimation bias. Experiments across a range of multi-turn agentic benchmarks show consistent gains over GRPO and superior performance compared to recent agentic RL algorithms. Code will be available at https://github.com/xcyao00/GRAFT.