Score
Designs, builds, and applies systematic methods and tools to identify, localize, and attribute the underlying causes of incidents, failures, or bugs across a system; this includes diagnostic procedures, inference and localization techniques, and automated troubleshooting pipelines. Develops and validates corrective actions or mitigations and integrates root-cause findings into debugging, incident reports, and automation to prevent recurrence.
Existing root cause analysis (RCA) research lacks a goal-oriented, systematic taxonomy, leading to task ambiguity and hindered progress assessment. Method: This paper proposes the first RCA classification framework centered on fundamental objectives—departing from conventional data-type–based taxonomies—and systematically categorizes 135 studies (2014–2025) according to core goals such as fault localization and defect remediation. Guided by a systematic literature review, we construct a multi-level RCA objective hierarchy that characterizes the state of the art, recurrent challenges, and critical technical gaps per task. Contribution/Results: We present the first RCA objective-method mapping atlas tailored to cloud service scenarios, establishing a theoretical foundation for academic research and a practical technology roadmap for industrial deployment.
This paper addresses the dual challenges of low fault localization accuracy and weak root-cause interpretability in software debugging. To this end, we propose an interpretable diagnosis method based on multi-execution feature fusion. Through empirical analysis of 310 real-world defects, we first establish—systematically and for the first time—that scalar pairs constitute the strongest failure-correlated features. Building upon this insight, we design a joint modeling framework that integrates 17 fine-grained execution features, including variable values, branch conditions, and definition-use chains. We further develop a feature-importance-driven interpretable decision tree model that automatically generates human-readable diagnostic rules. Evaluation across 20 open-source projects demonstrates that our approach significantly improves both fault localization accuracy and root-cause identification depth, substantially reducing developer debugging time. The method achieves a favorable balance between high precision and strong interpretability.
研究通过轨迹级分析方法评估LLM代理在微服务根因分析中的表现,提出DiagGuard框架以提高诊断准确性。
Bug fixing is a complex and time-consuming task in software development. Bug localization research tends to focus on the accuracy of automated tools that suggest source code files for developers to look at. However, little is known about how developers use these tools in practice. This paper reports on an ongoing qualitative user study. Eleven participants worked through four realistic bug localization tasks in a controlled environment and were given varying levels of support information offered by a specialized tool. Participants were asked to think aloud in a semi-structured interview session. The preliminary findings provide insight into three aspects of practice: how developers interact with tools, the role social and contextual information plays, and problem solving. The study demonstrates that bug localization is complex and suggests that the adoption of effective tools depends on more than their accuracy.
Systematic root-cause diagnosis for failures in platformized multi-agent systems remains underexplored. Method: We introduce AgentFail—a first-of-its-kind, fine-grained annotated failure log dataset (307 samples)—and propose the first taxonomy for failure root-cause classification in this domain. Leveraging this taxonomy, we design a classification-guided large language model prompting framework that integrates counterfactual reasoning and human verification to ensure annotation reliability, and release a reproducible automated diagnosis benchmark. Results: Experiments reveal that state-of-the-art methods achieve only 33.6% accuracy, underscoring the task’s substantial difficulty. Our work provides empirical evidence and practical guidelines for enhancing the robustness of multi-agent system design, establishing foundational resources and evaluation protocols for future research.
Inaccurate root cause analysis (RCA) in distributed systems stems from incomplete fault reports on platforms like GitHub and JIRA, lacking sufficient code-level diagnostic context. Method: This paper proposes a code-knowledge-enhanced RCA framework that automatically extracts relevant code snippets via static analysis, reconstructs exception propagation paths and call contexts, and dynamically injects code-level diagnostic signals—including call chains, function signatures, and exception-handling logic—into large language model (LLM) inference to align problem descriptions with source-code semantics and enable collaborative reasoning. The method integrates execution-path reconstruction, multi-example prompt engineering, and generative LLM inference, ensuring cross-system and cross-model generalizability. Results: Evaluated on five real-world distributed-system datasets, the framework improves root cause localization accuracy by 28.3% and root cause summary quality by 22.0%, while maintaining robust performance across multiple mainstream LLMs.
Industrial root cause diagnosis typically relies on manual hypotheses and extensive fault labels, yet existing data-driven methods suffer from poor interpretability and limited generalization. This work proposes AgentRCA, a novel framework that achieves zero-shot, label-free root cause diagnosis for the first time. By integrating data-driven digital twins with tool-augmented large language models, AgentRCA adopts a hypothesis-driven approach to iteratively gather statistical evidence, evaluate competing hypotheses, and construct transparent reasoning chains that explicitly link observed symptoms to underlying physical faults. Evaluated on real-world multiphase flow facilities and large-scale chemical plants, the framework matches the diagnostic performance of fully supervised baselines while offering high interpretability.
研究利用LLM系统辅助工业夜间测试失败的根因分析,通过单代理和多代理配置实现RCA流程,评估表明单代理系统在成本和速度上更优。
This study addresses the lack of empirical guidance on tool design and composition for large language model (LLM) agents in microservice root cause analysis (RCA) by constructing the first systematic empirical benchmark dedicated to agentic RCA tool abstraction and composition. We propose a hierarchical tool architecture spanning levels L0 through L3 and conduct multi-model comparative experiments alongside trajectory analysis to quantitatively evaluate how different tool configurations affect diagnostic performance. Results demonstrate that higher-level tools (L3) halve fault localization time while improving fault type identification, revealing inherent accuracy-efficiency trade-offs across tool hierarchy levels. These findings provide data-driven decision-making foundations for agent tool selection and design in automated microservice diagnostics.
Existing root cause analysis (RCA) research is constrained by Top@k metrics, making it difficult to distinguish between failures originating in the retrieval and reranking stages. This work proposes a retrieval-reranking decoupled framework that constructs a two-stage pipeline comprising a multi-signal fusion retriever and a large language model-based reranker. The proposed approach achieves high-precision root cause localization without requiring causal graphs or annotated data. Experimental results demonstrate that this method comprehensively outperforms the strongest baselines across six benchmarks, improving Top@1 accuracy by up to 18 percentage points.
While microservice failures are readily detectable, root cause analysis remains inefficient due to alarm flooding and the absence of structured memory capturing system dependencies and historical behaviors. This work proposes a topology-aware, operation-memory-driven multi-agent architecture that decouples root cause inference from explanation for the first time: the former relies on deterministic computation using a learned dependency graph and temporal anomaly thresholds, while the latter leverages a large language model to generate interpretable recommendations grounded in structured evidence. A novel four-layer operational memory mechanism enables traceable and reusable autonomous operations. Evaluated on an e-commerce benchmark platform with eight types of injected faults, the approach successfully reproduces and resolves two real-world cascading failures, significantly improving diagnostic accuracy and efficiency.