🤖 AI Summary
This study addresses the credit assignment misalignment in terminal agent training, where existing methods fail to track inter-command read-write dependencies. To overcome this limitation, this work proposes Dependency-aware Group Policy Optimization (DepGPO), which is the first approach to leverage explicit inter-command read-write dependencies for guiding fine-grained credit assignment in reinforcement learning. Specifically, DepGPO parses execution trajectories, constructs command dependency graphs, and backtraces validator resources to precisely identify critical steps and reallocate trajectory-level advantages accordingly. Experimental results demonstrate that the proposed method significantly improves both performance and training stability on complex terminal tasks by effectively enhancing the learning of critical steps.
📝 Abstract
Terminal-using agents benefit from reinforcement learning (RL) in coding, debugging, and other multi-step terminal tasks. In these tasks, later commands often depend on information or intermediate results produced by earlier commands. However, existing trajectory-level and step-level credit assignment methods do not explicitly trace the read-write dependencies through which commands affect the final outcome. Consequently, training signals could still be assigned to irrelevant operations, weakening learning from relevant steps. In this paper, we propose Dependency-Aware Group Policy Optimization (DepGPO), which uses execution dependencies between commands to guide credit assignment for terminal agents. Specifically, we construct a command dependency graph from execution traces and trace backward from the resources inspected by the task verifier. We then assign credit to relevant writes and their supporting reads along these paths, and use it to redistribute trajectory advantages across steps. Extensive comparative experiments and ablation studies demonstrate that DepGPO improves task performance and training stability on complex terminal tasks.