Argus: Agentic, Reference-Calibrated, Tree-Guided, System-Software-Level Bottleneck Localization

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of localizing operating system performance bottlenecks and the reliance of existing tools on manual effort, which often introduces measurement interference. To overcome these limitations, this work proposes an automated kernel diagnosis method driven by large language model (LLM) agents. The approach enables LLM agents to automatically generate instrumentation code while introducing an idle-system reference calibration mechanism and a tree-structured execution path data representation. These innovations facilitate autonomous reasoning from subsystems down to specific code paths, achieving deep and precise diagnosis. Experimental results demonstrate that, compared to the strongest baseline, the proposed method reduces the deep-level misdiagnosis rate by 19-fold with a diagnostic latency of approximately 31 seconds. Furthermore, it successfully identifies critical performance bottlenecks within memory management, validating its effectiveness for practical kernel optimization.
📝 Abstract
Operating system (OS) code can account for a substantial share of CPU execution time. First, as application logic is offloaded to heterogeneous accelerators (e.g., GPUs), the CPU increasingly acts as an orchestrator, spending cycles in driver calls, data movement, and synchronization rather than in application code. Second, workloads such as serverless functions frequently invoke OS services. At the same time, the OS is a complex codebase spanning many subsystems (e.g., memory management, networking), making it hard to localize the specific code path responsible for a slowdown. Existing profilers expose measurements that require interpretation(e.g., perf and Intel VTune) or can perturb short operations when extensively instrumented (e.g., ftrace). Diagnosing OS bottlenecks can therefore require repeated kernel instrumentation and manual interpretation. We introduce Argus, an agentic LLM-based profiler that produces instrumentation code and autonomously reasons over potential OS-level bottlenecks. Argus integrates two key mechanisms: (i) a calibration methodology that involves collecting a measurement from an idle system and using it as a reference point to discover potential bottlenecks, and (ii) a tree-based data structure that represents the different OS execution paths, improving the agent's bottleneck localization accuracy. Argus aims to identify a specific kernel code path rather than stop at a subsystem-level diagnosis. In two case studies, we employ Argus to autonomously discover bottlenecks present in the memory management subsystem caused by (i) a THP aggressor co-running with other applications, and (ii) applications that incur different types of page faults. Argus produces 19 times fewer incorrect deep-path diagnoses than the strongest evaluated LLM-based baseline, which lacks reference calibration, while preserving low time-to-diagnosis (approximately 31 s)
Problem

Research questions and friction points this paper is trying to address.

operating system bottleneck
performance profiling
bottleneck localization
kernel code path
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic LLM Profiler
Reference Calibration
Tree-Guided Localization
OS Bottleneck Diagnosis
Kernel Instrumentation
🔎 Similar Papers
No similar papers found.
V
Vlad-Petru Nitu
ETH Zürich
H
Harsh Songara
ETH Zürich
K
Konstantinos Sgouras
ETH Zürich
S
Spiros Galanopoulos
ETH Zürich
Konstantinos Kanellopoulos
Konstantinos Kanellopoulos
ETH Zurich
Onur Mutlu
Onur Mutlu
ETH Zürich and Carnegie Mellon University
Computer ArchitectureMemory SystemsEnergy EfficiencyHardware SecurityGenome Analysis