Credit Where It Matters: Dependency-Aware Policy Optimization for Terminal Agents

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the credit assignment misalignment in terminal agent training, where existing methods fail to track inter-command read-write dependencies. To overcome this limitation, this work proposes Dependency-aware Group Policy Optimization (DepGPO), which is the first approach to leverage explicit inter-command read-write dependencies for guiding fine-grained credit assignment in reinforcement learning. Specifically, DepGPO parses execution trajectories, constructs command dependency graphs, and backtraces validator resources to precisely identify critical steps and reallocate trajectory-level advantages accordingly. Experimental results demonstrate that the proposed method significantly improves both performance and training stability on complex terminal tasks by effectively enhancing the learning of critical steps.
📝 Abstract
Terminal-using agents benefit from reinforcement learning (RL) in coding, debugging, and other multi-step terminal tasks. In these tasks, later commands often depend on information or intermediate results produced by earlier commands. However, existing trajectory-level and step-level credit assignment methods do not explicitly trace the read-write dependencies through which commands affect the final outcome. Consequently, training signals could still be assigned to irrelevant operations, weakening learning from relevant steps. In this paper, we propose Dependency-Aware Group Policy Optimization (DepGPO), which uses execution dependencies between commands to guide credit assignment for terminal agents. Specifically, we construct a command dependency graph from execution traces and trace backward from the resources inspected by the task verifier. We then assign credit to relevant writes and their supporting reads along these paths, and use it to redistribute trajectory advantages across steps. Extensive comparative experiments and ablation studies demonstrate that DepGPO improves task performance and training stability on complex terminal tasks.
Problem

Research questions and friction points this paper is trying to address.

credit assignment
terminal agents
reinforcement learning
read-write dependencies
multi-step tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dependency-Aware Policy Optimization
Credit Assignment
Command Dependency Graph
Terminal Agents
Reinforcement Learning
Y
Yu Li
School of Computer Science and Engineering, Southeast University, Nanjing, China
G
Guangfeng Cai
School of Computer Science and Engineering, Southeast University, Nanjing, China
Long-Fei Li
Long-Fei Li
Huawei Noah’s Ark Lab, China
S
Shuo Han
Huawei Noah’s Ark Lab, China
S
Shengtian Yang
School of Computer Science and Engineering, Southeast University, Nanjing, China
H
Han Luo
School of Computer Science and Engineering, Southeast University, Nanjing, China
K
Kaibing Yang
School of Computer Science and Engineering, Southeast University, Nanjing, China
Lei Feng
Lei Feng
Professor, Southeast University
Machine LearningData ScienceStatistics