DAAF: From Failure Localization to Editable System Assets in LLM Agents

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the decision gap between fault localization and component repair in LLM agents by proposing a Detection-Aware Attribution Framework (DAAF) designed to precisely identify versioned components requiring modification for improved task performance. DAAF amortizes intervention evidence into deployment diagnostics, integrates sparse noisy signals to determine intervention timing, and learns component replacement effects through controlled replay. Its core innovation lies in achieving attribute-level fault attribution without counterfactual replay or task rewards, deriving repair targets solely from observed executions. Evaluated on the Telecom benchmark, DAAF attains an 80.72% attribute hit rate and a 71.93% task success rate, recovering 62.65% of failed executions with a regression rate of only 3.23%.
📝 Abstract
Deployed LLM agents increasingly rely on persistent, versioned system assets such as routing rules, knowledge segments, prompt instructions, and reusable skills. Failure-localization methods can identify where an error manifests in an agent or execution trace, but repair requires a different decision: which editable system asset should be changed, and is that change expected to improve the task outcome? We study this gap through component-attribute failure attribution, where diagnosis targets versioned, addressable items rather than execution locations. We propose the Detection-Aware Attribution Framework (DAAF), which learns the effects of valid attribute replacements and amortizes this intervention evidence into deployment-time diagnosis. DAAF combines sparse and noisy failure signals to decide whether intervention is warranted, learns component-type-conditioned replacement effects from controlled replays evaluated by executable task outcomes, and shares supervision across requests with compatible intervention responses. At diagnosis time, DAAF uses only the observed execution, registered candidates, and available failure signals; it requires neither counterfactual replay nor task reward and returns no_change, a repair target, or an unresolved decision when evidence is insufficient. On held-out tau^2-bench Telecom tasks, DAAF achieves 80.72% attribute Hit@1, recovers 62.65% of failed executions while limiting clean-task regression to 3.23%, and reaches 71.93% overall task success. These results show that intervention-grounded attribute attribution can connect failure localization to executable system repair.
Problem

Research questions and friction points this paper is trying to address.

LLM agents
failure attribution
system repair
failure localization
editable system assets
Innovation

Methods, ideas, or system contributions that make the work stand out.

Failure Attribution
LLM Agents
System Repair
Intervention-grounded Diagnosis
DAAF
🔎 Similar Papers
No similar papers found.
X
Xiaoyang Yuan
Tencent
Q
Qi Liu
Tencent
Y
Yubin Ruan
Tencent
Xinyi Mou
Xinyi Mou
Fudan University
NLPLarge Language ModelsSocial Simulation
Z
Zhuomeng Zhang
Tencent
W
Wenjin Wang
Tencent
H
Hanying Jiao
Tencent
D
Di Wu
Tencent
Mingye Xu
Mingye Xu
Tencent
Yi Bin
Yi Bin
National University of Singapore
multimediavision and languagedeep learning
K
Ke Feng
Tencent
Z
Zixun Sun
Tencent