🤖 AI Summary
This study addresses the challenge of simultaneously achieving accuracy, cost-efficiency, and interpretability in large-scale internet service troubleshooting by proposing the E4 system. Its core innovation lies in a "constrained creativity" paradigm: rather than permitting arbitrary code generation by large language models (LLMs), it introduces a domain-specific language (DSL) equipped with high-level operators to guide LLM agents in automatically synthesizing loop-free dataflow programs. This approach ensures that outputs remain accurate, verifiable, and interpretable while substantially reducing operational overhead. Experimental results demonstrate that under mixed workloads, E4 improves diagnostic accuracy by 62% over state-of-the-art methods and reduces troubleshooting costs by up to 12×.
📝 Abstract
System administrators of Internet-scale services need to resolve failure incidents to maintain reliability of such services. Ideally, we want a troubleshooting system to be: (1) expressive to known and unknown incidents with high accuracy; (2) cost efficient at scale; (3) explainable to provide actionable insights operators can act on; and (4) entail low effort from the operators. Unfortunately, most existing systems, including emerging LLM-assisted agentic workflows and structured frameworks for authoring diverse RCA algorithms fall short of achieving all four requirements. We present E4, a novel agentic system for troubleshooting for Internet-scale services. E4 embodies the paradigm of constrained creativity that combines the best of LLM-assisted automation and exploration with the explainability and efficiency of a structured approach. Instead of allowing an LLM agent to write arbitrary code or generate arbitrary responses, we provide the agent a restricted DSL to generate its response via simple loop-free data flow programs. This DSL, equipped with high level operators for troubleshooting, makes E4's output accurate, verifiable and explainable. On a mix of synthetic and real-world workloads, E4 achieves up to 62% better accuracy compared to state-of-the-art solutions, while providing more explainable responses at up to 12x reduced cost.