SRE-Marathon: A Continuous, Change-Driven Benchmark for Autonomous Site Reliability Agents

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing Site Reliability Engineering (SRE) benchmarks, which predominantly evaluate isolated incidents and fail to capture real-world production complexities such as noisy alerts, overlapping failures, and change-driven, long-horizon operations. To bridge this gap, we introduce the first long-horizon, change-driven SRE evaluation paradigm, establishing a continuous operations benchmark tailored for autonomous agents. Leveraging a dual-zone Kubernetes environment with injected concurrent failures, our framework integrates CI/CD pipelines, sealed record bundles, and deterministic offline scoring mechanisms to provide cumulative alerts and persistent workspaces that faithfully replicate production complexity. Experimental results demonstrate that the best-performing method achieves only 41.3 points, revealing that while current agents can effectively correlate and localize faults, executing remediation during active failure windows remains a critical bottleneck.
📝 Abstract
Benchmarks for site reliability engineering (SRE) agents are typically episodic: one fault is injected, the agent receives an incident task, and its response is scored. Production operation is not. Incidents surface through noisy alerts, overlap in time, and often originate from code or configuration changes. We present SRE-Marathon, a benchmark for long-horizon, continuous SRE operation. An agent is invoked at a fixed cadence with cumulative alert history and a persistent workspace while operating a live two-zone Kubernetes deployment as a fault orchestrator injects overlapping faults according to a seeded, production-calibrated schedule. Curated code and configuration changes deployed through the same build pipeline available for repair. Each run is recorded into a sealed bundle and scored offline: Marathon-Score credits each injected fault for ordered progress through correlation, localization, and repair, with all metrics computed deterministically from recorded system evidence. Across three applications and about sixty faults per run, the best of 10 methods reaches only 41.3 out of 100. Agents often correlate and localize faults, but almost never complete repairs while the faults remain active.
Problem

Research questions and friction points this paper is trying to address.

Site Reliability Engineering
Autonomous Agents
Benchmark
Fault Management
Continuous Operation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Site Reliability Engineering
Continuous Benchmark
Autonomous Agents
Fault Orchestration
Kubernetes
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Yifang Tian
Yifang Tian
University of Toronto
LLMgraph neural networkroot cause analysisAIOpstransfer learning
Y
Yingjian Bai
University of Toronto
Y
Yifeng He
University of Toronto
Z
Zichun Chong
University of Toronto
Y
Yuanchen Gao
The Hong Kong University of Science and Technology
Y
Yiran Li
University of Toronto
Hans-Arno Jacobsen
Hans-Arno Jacobsen
Professor of Computer Engineering and Computer Science
data managementmiddlewaredistributed systemsevent processingblockchains