UndoBench: Separating Task Competence from Recovery Capability in Tool-Using AI Agents

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing AI agent benchmarks that conflate task planning with fault recovery, thereby hindering the assessment of operational safety in enterprise environments. To this end, we propose UndoBench, a benchmark introducing a novel evaluation paradigm that decouples task completion from fault recovery. Methodologically, counterfactual paired trials are employed to disentangle task competence from recovery capability, while line-level effect histories and environment state oracles are incorporated for precise verification across diverse domain workflows and multiple recovery strategies. Experimental results demonstrate that although agents achieve a nominal success rate of 83.54%, their conditional recovery success rate drops to merely 46.72%. This discrepancy reveals critical safety vulnerabilities, indicating that naive retry mechanisms can readily trigger repeated external side effects.
📝 Abstract
Tool-using AI agents are increasingly deployed across enterprise software systems, yet widely used benchmarks primarily evaluate nominal task completion, conflating baseline planning competence with operational fault recovery. We introduce UndoBench, a benchmark spanning 36 base workflows and 36 fault scenarios across 8 enterprise domains, decoupling task competence from recovery capability via counterfactual paired trials under identical seeds alongside wire-level effect-history and environment-state oracles. On 12 held-out TEST workflows across two open-weight models, two frameworks, and three recovery paradigms (5,760 executions / 2,880 paired trials) in the frozen lost-acknowledgment study, nominal competence reached 83.54% while conditional recovery success rate (CRSR) fell to 46.72%, with naive retry producing duplicate external effects in 53.33% of trials. Extensions to commercial API models reproduced this competence-recovery separation. Evaluations across complementary execution boundaries show that recovery is phase-dependent: before mutation, methods perform similarly without duplicate effects among capable trials; during partial mutation, naive retry, per-call idempotency, and zero-privilege journaling collapse on the evaluated composite workflows; after commit but before acknowledgment, verification and server-side idempotency substantially improve safety. These findings demonstrate that evaluating nominal completion alone masks critical, phase-dependent recovery vulnerabilities in autonomous agents.
Problem

Research questions and friction points this paper is trying to address.

AI agents
fault recovery
benchmark evaluation
tool use
idempotency
Innovation

Methods, ideas, or system contributions that make the work stand out.

UndoBench
fault recovery capability
counterfactual paired trials
phase-dependent recovery
tool-using AI agents
D
Dolly Sah
Independent Researcher
T
Tanmay Sah
Independent Researcher
H
Harshul Jain
Independent Researcher
T
Tanya Sah
Independent Researcher