🤖 AI Summary
This study addresses the limitation of existing Site Reliability Engineering (SRE) benchmarks, which predominantly evaluate isolated incidents and fail to capture real-world production complexities such as noisy alerts, overlapping failures, and change-driven, long-horizon operations. To bridge this gap, we introduce the first long-horizon, change-driven SRE evaluation paradigm, establishing a continuous operations benchmark tailored for autonomous agents. Leveraging a dual-zone Kubernetes environment with injected concurrent failures, our framework integrates CI/CD pipelines, sealed record bundles, and deterministic offline scoring mechanisms to provide cumulative alerts and persistent workspaces that faithfully replicate production complexity. Experimental results demonstrate that the best-performing method achieves only 41.3 points, revealing that while current agents can effectively correlate and localize faults, executing remediation during active failure windows remains a critical bottleneck.
📝 Abstract
Benchmarks for site reliability engineering (SRE) agents are typically episodic: one fault is injected, the agent receives an incident task, and its response is scored. Production operation is not. Incidents surface through noisy alerts, overlap in time, and often originate from code or configuration changes. We present SRE-Marathon, a benchmark for long-horizon, continuous SRE operation. An agent is invoked at a fixed cadence with cumulative alert history and a persistent workspace while operating a live two-zone Kubernetes deployment as a fault orchestrator injects overlapping faults according to a seeded, production-calibrated schedule. Curated code and configuration changes deployed through the same build pipeline available for repair. Each run is recorded into a sealed bundle and scored offline: Marathon-Score credits each injected fault for ordered progress through correlation, localization, and repair, with all metrics computed deterministically from recorded system evidence. Across three applications and about sixty faults per run, the best of 10 methods reaches only 41.3 out of 100. Agents often correlate and localize faults, but almost never complete repairs while the faults remain active.