From Anomalies to Failures: Constructing Causal Error Graphs for Agentic Trace Diagnosis

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of conflating anomalies with failures and the absence of causal propagation in fault diagnosis over long trajectories of LLM agents. To this end, it proposes the CEG-Agent framework, which first establishes an explicit taxonomy distinguishing anomalies, errors, and failures. It further introduces a unified, typed causal error graph representation and integrates a tool-augmented architecture with a multi-agent adversarial arbitration mechanism to achieve precise fault attribution. Experimental results demonstrate that the proposed method attains state-of-the-art performance on the CEG-Bench benchmark across both semantic and structural evaluations, with its annotations exhibiting strong alignment with expert consensus.
📝 Abstract
LLM-driven agents are increasingly deployed in complex applications, where long agentic traces make failures difficult to diagnose. Existing trace diagnosis methods often conflate anomalies, errors, and failures, making diagnostic targets ambiguous; they also lack structured modeling of how causally relevant errors propagate and amplify into final task failures, resulting in unreliable failure attribution. To address these problems, we propose CEG-Agent, a tool-augmented agentic framework for causal diagnosis of agentic traces. Specifically, CEG-Agent introduces an explicit taxonomy of anomalies, errors, and failures, and constructs Causal Error Graphs (CEGs), a unified typed representation that links execution events, diagnostic nodes, and failure outcomes through causal relations. To evaluate causal trace diagnosis, we further construct CEG-Bench, a fully agent-annotated benchmark with high-confidence, consensus-derived CEG annotations obtained through an Adversarial Agentic Adjudication Protocol (AAAP). We validate the resulting annotations against an expert-curated human gold set, which shows close agreement with the automatic annotations. Experiments on CEG-Bench demonstrate that CEG-Agent achieves state-of-the-art performance under both semantically relaxed and structurally exact evaluation criteria. Our code is publicly available.
Problem

Research questions and friction points this paper is trying to address.

Agentic trace diagnosis
Causal error graphs
Failure attribution
LLM-driven agents
Anomaly taxonomy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Causal Error Graphs
Agentic Trace Diagnosis
CEG-Agent
Adversarial Agentic Adjudication Protocol
CEG-Bench
S
Shu-Xun Yang
Beijing Institute of Technology, Beijing, China; Zhipu AI, Beijing, China
Y
Yidong Wang
Zhipu AI, Beijing, China
Zhuoer Feng
Zhuoer Feng
Undergraduate, Tsinghua University
artificial intelligencedeep learning
Bosi Wen
Bosi Wen
Tsinghua University
Natural Language Processing
J
Jiayi Gui
Zhipu AI, Beijing, China
D
Dayong Yang
Zhipu AI, Beijing, China
W
Wenbo Yu
Zhipu AI, Beijing, China
H
Haoke Zhang
Zhipu AI, Beijing, China
Jie Tang
Jie Tang
UW Madison
Computed Tomography
Cunxiang Wang
Cunxiang Wang
Tsinghua University; ZhipuAI
Large Language ModelsLLM EvaluationLLM Post-training