Incident-Arena: Getting agents to the last nine of reliability

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing AI coding agents in production incident response, where benchmarks are often unrealistic and validation methods are overly simplistic. To this end, we construct a benchmark comprising 20 fault injection tasks grounded in real-world open-source software and Kubernetes clusters. We further propose a novel functional verification framework that integrates dynamic workloads, authentic fault injection, and system-level metric stability, thereby overcoming the constraints of traditional static checks. Experimental results demonstrate that state-of-the-art large language models achieve an average score below 64.3%, revealing critical bottlenecks such as diagnostic bias, incomplete remediation, and safety regressions during extended reasoning processes. This work establishes a new paradigm for evaluating and enhancing the reliability of AI coding agents in complex operational environments.
📝 Abstract
AI coding agents are ubiquitous in engineering workflows amongst industry and academia. Yet, despite their use in app coding, relatively less attention has been paid to their ability to execute on production incident response. This emerging field, termed agentic site-reliability-engineering (SRE) contains benchmarks limited by (1) unrealistic environments, typically toy repositories (2) non-standard framework implementations and (3) simple static verifiers. We introduce Incident-Arena, a human-built benchmark of 20 carefully selected tasks grounded in real-world deployed open source software. Each task deploys a production application to an ephemeral Kubernetes cluster, injecting a fault from the config layer through underlying images, and a sustained load profile given the task requirements. We also present a novel verification method, going beyond static checks to functional verifiers, holding systems level metrics stable, while ensuring repairs are done safely. Agent trials run an average of 2.81M tokens and 41 turns, going beyond existing benchmarks, demonstrating agentic long horizon reasoning. Across 20 tasks and 3 application substrates, frontier models score below 64.3%, with failures extending from diagnosis/localization errors, through incomplete repairs and unsafe regressions.
Problem

Research questions and friction points this paper is trying to address.

AI coding agents
Site Reliability Engineering (SRE)
benchmark evaluation
incident response
long horizon reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic SRE
Incident-Arena
Functional Verification
Kubernetes Fault Injection
Long Horizon Reasoning
💼 Related Jobs
No related jobs found.
A
Andre Fu
Abundant AI
M
Malik Drabla
Adrenaline AI
L
Leon Liu
Carnegie Mellon University
M
Meji Abidoye
Abundant AI
Marek Šuppa
Marek Šuppa
Comenius University in Bratislava
Natural Language ProcessingComputer VisionMachine Learning
L
Lata Mishra
Abundant AI
A
Adnan El Assadi
Massachusetts General Hospital
Yiyuan Li
Yiyuan Li
University of North Carolina at Chapel Hill
Natural Language ProcessingComputational Linguistics