SSCBench: Evaluating the Evidential Validity of Fault-Injection Tests for Tool-Using LLM Agents

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of interpreting evidence validity in fault injection evaluations for tool-use LLM agents, where observation visibility varies. For the first time, it incorporates the temporal conditions of refuting evidence into the evaluation framework. By constructing the SSCBench benchmark and a standardized measurement protocol, combined with automated trajectory analysis and multi-environment agent experiments, this work systematically quantifies how the timely visibility of refuting evidence during execution influences fault adoption behavior. The findings reveal significant discrepancies between aggregate adoption rates and individual temporal diagnostics, alongside sparse support phenomena. These results demonstrate that fault adoption must be interpreted in conjunction with specific evidential conditions, cautioning against misjudging agent reliability based solely on aggregate metrics.
📝 Abstract
Fault injection is increasingly used to evaluate the reliability of tool-using LLM agents. However, there has been limited study of how fault-adoption results should be interpreted when the agent itself determines which authoritative observations become visible during execution. In this paper, we present a systematic study of this evidential validity problem in agent fault-injection evaluation. We develop a measurement protocol that specifies what observations can refute an injected assertion, determines whether they can become visible before the affected fact is first used, and records whether the evaluated execution actually realizes this condition. We construct SSCBench as an instantiation of the protocol and evaluate four fault operators and five agent configurations over 1,191 faulted executions in two $τ$-bench environments. Our experiments show that an admitted fault case and agent configuration can realize substantially different evidential conditions across executions, and that aggregate adoption can remain well defined even when the population supporting a timely-counterevidence claim is sparse or absent. For example, among 44 adopted runs in which counterevidence eventually became visible, only 17 received it before first use, while 27 received it afterward. We also find that first-error timing and later stance revision need not coincide, and that automated trajectory analysis can recover adoption without reliably recovering the first faulty-reliance event needed for temporal diagnosis. We argue that the evidential condition realized by an execution and the population supporting a claim-specific interpretation are part of fault-injection evaluation itself and should be reported before adoption is interpreted as failure under pre-use counterevidence.
Problem

Research questions and friction points this paper is trying to address.

fault injection
tool-using LLM agents
evidential validity
reliability evaluation
counter-evidence visibility
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fault Injection
Evidential Validity
LLM Agents
Benchmark
Trajectory Analysis
X
Xincheng He
School of Artificial Intelligence and Computer Science, Jiangnan University, China
W
Wanli Dong
School of Artificial Intelligence and Computer Science, Jiangnan University, China
Z
Zhaoqiang Guo
College of Computer Science and Technology, Zhejiang University, China
Yan Liu
Yan Liu
Indiana University, California Institute of Technology, Washington University
Wavefront shapingAdaptive opticsOptical coherence tomographyOcular imagingPhotoacoustics
L
Lei Xu
State Key Laboratory for Novel Software Technology, Nanjing University, China