DeFA: Dependency-Guided Failure Attribution for LLM Agents

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the attribution challenge in LLM-based multi-agent systems caused by the spatiotemporal decoupling of errors and their consequences, proposing a novel dependency-guided diagnostic framework. The method integrates protocol and semantic dependencies to construct event dependency graphs and fault propagation graphs. Furthermore, it introduces a trajectory segmentation and aggregation mechanism that combines local details with global summaries to precisely identify decisive errors and responsible agents, thereby enabling closed-loop optimization from fault attribution to skill evolution. Experimental results demonstrate that the proposed framework achieves state-of-the-art accuracy in identifying responsible agents and erroneous steps on the Who and When benchmark, while improving downstream task performance by 6–15 percentage points.
📝 Abstract
Errors in LLM agent executions and their visible consequences can be separated by many steps, making decisive-error localization a matter of understanding both step content and step dependencies. We introduce DeFA, a dependency-guided framework for agent failure attribution. DeFA first combines protocol relations and semantic dependencies into an event dependency graph spanning the trajectory. It then identifies events that may violate task requirements and traces their sources and subsequent effects to construct a failure propagation graph. Finally, DeFA uses step evidence and the steps' roles in failure propagation to identify the decisive error, responsible agent, and error category. To support long trajectories, DeFA partitions executions into segments and combines the current segment's detailed content with summaries of the other segments, giving local diagnosis access to global execution context. Across Who and When and the Who and When Pro text subset, DeFA achieves the highest responsible-agent and exact step accuracy with all evaluated backbones, and the highest failure-mode accuracy among taxonomy-aligned methods on Pro. Further experiments on image and video trajectories demonstrate its applicability to multimodal failure attribution. Ablations support the contributions of segmentation, the event dependency graph, and the failure propagation graph. Using DeFA's diagnostic feedback for skill evolution in Trace2Skill improves downstream task accuracy by 6-15 percentage points over the native pipeline, showing that the diagnoses can also support agent improvement on subsequent tasks.
Problem

Research questions and friction points this paper is trying to address.

Failure Attribution
LLM Agents
Error Localization
Step Dependencies
Decisive Error
Innovation

Methods, ideas, or system contributions that make the work stand out.

Failure Attribution
Dependency Graph
LLM Agents
Long Trajectory Segmentation
Multimodal
🔎 Similar Papers
No similar papers found.
B
Bo Deng
Qwen DianJin Team, Alibaba Cloud Computing
X
Xinlei Zheng
Beihang University
Y
Yi Wei
Beihang University
K
Kang Zhou
Qwen DianJin Team, Alibaba Cloud Computing
Chongyang Tao
Chongyang Tao
Associate Professor of Computer Science, Beihang University
Natural Language ProcessingDialogue SystemsInformation RetrievalData Intelligence
R
Renzhao Liang
Beihang University
X
Xuanren Chen
Beihang University
Lifan Guo
Lifan Guo
Researcher Drexel University
Machine Learning
C
Chi Zhang
Qwen DianJin Team, Alibaba Cloud Computing