TRACE: Diagnosing Verifier Brittleness in Agentic Evaluation
In the evaluation of large language model (LLM) agents, fluctuations in verifier scores frequently conflate genuine capability changes with assessment bias. This work proposes the TRACE protocol, which systematically modifies individual evaluation components and conducts paired comparative runs to transform score variations into a testable causal diagnostic process, thereby precisely disentangling agent behavioral shifts from scoring rule artifacts. Through experiments involving synthetic tasks, public benchmarks, and repeated multi-agent trials, this study reveals that tool renaming induces spurious score degradations and high variance. The results demonstrate that TRACE effectively identifies measurement errors, offering a reliable attribution analysis framework for robust agent evaluation.