🤖 AI Summary
Inaccurate root cause analysis (RCA) in distributed systems stems from incomplete fault reports on platforms like GitHub and JIRA, lacking sufficient code-level diagnostic context. Method: This paper proposes a code-knowledge-enhanced RCA framework that automatically extracts relevant code snippets via static analysis, reconstructs exception propagation paths and call contexts, and dynamically injects code-level diagnostic signals—including call chains, function signatures, and exception-handling logic—into large language model (LLM) inference to align problem descriptions with source-code semantics and enable collaborative reasoning. The method integrates execution-path reconstruction, multi-example prompt engineering, and generative LLM inference, ensuring cross-system and cross-model generalizability. Results: Evaluated on five real-world distributed-system datasets, the framework improves root cause localization accuracy by 28.3% and root cause summary quality by 22.0%, while maintaining robust performance across multiple mainstream LLMs.
📝 Abstract
Runtime failures are commonplace in modern distributed systems. When such issues arise, users often turn to platforms such as Github or JIRA to report them and request assistance. Automatically identifying the root cause of these failures is critical for ensuring high reliability and availability. However, prevailing automatic root cause analysis (RCA) approaches rely significantly on comprehensive runtime monitoring data, which is often not fully available in issue platforms. Recent methods leverage large language models (LLMs) to analyze issue reports, but their effectiveness is limited by incomplete or ambiguous user-provided information. To obtain more accurate and comprehensive RCA results, the core idea of this work is to extract additional diagnostic clues from code to supplement data-limited issue reports. Specifically, we propose COCA, a code knowledge enhanced root cause analysis approach for issue reports. Based on the data within issue reports, COCA intelligently extracts relevant code snippets and reconstructs execution paths, providing a comprehensive execution context for further RCA. Subsequently, COCA constructs a prompt combining historical issue reports along with profiled code knowledge, enabling the LLMs to generate detailed root cause summaries and localize responsible components. Our evaluation on datasets from five real-world distributed systems demonstrates that COCA significantly outperforms existing methods, achieving a 28.3% improvement in root cause localization and a 22.0% improvement in root cause summarization. Furthermore, COCA's performance consistency across various LLMs underscores its robust generalizability.