perform root cause analysis

Designs, builds, and applies systematic methods and tools to identify, localize, and attribute the underlying causes of incidents, failures, or bugs across a system; this includes diagnostic procedures, inference and localization techniques, and automated troubleshooting pipelines. Develops and validates corrective actions or mitigations and integrates root-cause findings into debugging, incident reports, and automation to prevent recurrence.

performrootcauseanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
1.5
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$192K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

How Execution Features Relate to Failures: An Empirical Study and Diagnosis Approach

Feb 25, 2025
MS
Marius Smytzek
🏛️ CISPA Helmholtz Center for Information Security | Humboldt-Universität zu Berlin

This paper addresses the dual challenges of low fault localization accuracy and weak root-cause interpretability in software debugging. To this end, we propose an interpretable diagnosis method based on multi-execution feature fusion. Through empirical analysis of 310 real-world defects, we first establish—systematically and for the first time—that scalar pairs constitute the strongest failure-correlated features. Building upon this insight, we design a joint modeling framework that integrates 17 fine-grained execution features, including variable values, branch conditions, and definition-use chains. We further develop a feature-importance-driven interpretable decision tree model that automatically generates human-readable diagnostic rules. Evaluation across 20 open-source projects demonstrates that our approach significantly improves both fault localization accuracy and root-cause identification depth, substantially reducing developer debugging time. The method achieves a favorable balance between high precision and strong interpretability.

Analyzing diverse execution featuresDeveloping interpretable debugging diagnosesEnhancing fault localization accuracy

Bug fixing is a complex and time-consuming task in software development. Bug localization research tends to focus on the accuracy of automated tools that suggest source code files for developers to look at. However, little is known about how developers use these tools in practice. This paper reports on an ongoing qualitative user study. Eleven participants worked through four realistic bug localization tasks in a controlled environment and were given varying levels of support information offered by a specialized tool. Participants were asked to think aloud in a semi-structured interview session. The preliminary findings provide insight into three aspects of practice: how developers interact with tools, the role social and contextual information plays, and problem solving. The study demonstrates that bug localization is complex and suggests that the adoption of effective tools depends on more than their accuracy.

bug localizationdeveloper behaviourqualitative study

Diagnosing Failure Root Causes in Platform-Orchestrated Agentic Systems: Dataset, Taxonomy, and Benchmark

Sep 28, 2025
XM
Xuyan Ma
🏛️ University of Chinese Academy of Sciences | Singapore Management University | Chinese Academy of Sciences

Systematic root-cause diagnosis for failures in platformized multi-agent systems remains underexplored. Method: We introduce AgentFail—a first-of-its-kind, fine-grained annotated failure log dataset (307 samples)—and propose the first taxonomy for failure root-cause classification in this domain. Leveraging this taxonomy, we design a classification-guided large language model prompting framework that integrates counterfactual reasoning and human verification to ensure annotation reliability, and release a reproducible automated diagnosis benchmark. Results: Experiments reveal that state-of-the-art methods achieve only 33.6% accuracy, underscoring the task’s substantial difficulty. Our work provides empirical evidence and practical guidelines for enhancing the robustness of multi-agent system design, establishing foundational resources and evaluation protocols for future research.

Addressing fragility of LLM-driven agent coordination systemsDeveloping taxonomy and benchmark for agentic system failure analysisIdentifying failure root causes in platform-orchestrated multi-agent systems

COCA: Generative Root Cause Analysis for Distributed Systems with Code Knowledge

Mar 29, 2025
YL
Yichen Li
🏛️ The Chinese University of Hong Kong | Sun Yat-sen University

Inaccurate root cause analysis (RCA) in distributed systems stems from incomplete fault reports on platforms like GitHub and JIRA, lacking sufficient code-level diagnostic context. Method: This paper proposes a code-knowledge-enhanced RCA framework that automatically extracts relevant code snippets via static analysis, reconstructs exception propagation paths and call contexts, and dynamically injects code-level diagnostic signals—including call chains, function signatures, and exception-handling logic—into large language model (LLM) inference to align problem descriptions with source-code semantics and enable collaborative reasoning. The method integrates execution-path reconstruction, multi-example prompt engineering, and generative LLM inference, ensuring cross-system and cross-model generalizability. Results: Evaluated on five real-world distributed-system datasets, the framework improves root cause localization accuracy by 28.3% and root cause summary quality by 22.0%, while maintaining robust performance across multiple mainstream LLMs.

Automatically identify root causes of runtime failures in distributed systemsEnhance RCA by extracting code clues from incomplete issue reportsImprove accuracy of root cause localization and summarization using LLMs

Latest Papers

What's happening recently
View more

Industrial root cause diagnosis typically relies on manual hypotheses and extensive fault labels, yet existing data-driven methods suffer from poor interpretability and limited generalization. This work proposes AgentRCA, a novel framework that achieves zero-shot, label-free root cause diagnosis for the first time. By integrating data-driven digital twins with tool-augmented large language models, AgentRCA adopts a hypothesis-driven approach to iteratively gather statistical evidence, evaluate competing hypotheses, and construct transparent reasoning chains that explicitly link observed symptoms to underlying physical faults. Evaluated on real-world multiphase flow facilities and large-scale chemical plants, the framework matches the diagnostic performance of fully supervised baselines while offering high interpretability.

anomaly diagnosisexplainable AIindustrial operation

This study addresses the lack of empirical guidance on tool design and composition for large language model (LLM) agents in microservice root cause analysis (RCA) by constructing the first systematic empirical benchmark dedicated to agentic RCA tool abstraction and composition. We propose a hierarchical tool architecture spanning levels L0 through L3 and conduct multi-model comparative experiments alongside trajectory analysis to quantitatively evaluate how different tool configurations affect diagnostic performance. Results demonstrate that higher-level tools (L3) halve fault localization time while improving fault type identification, revealing inherent accuracy-efficiency trade-offs across tool hierarchy levels. These findings provide data-driven decision-making foundations for agent tool selection and design in automated microservice diagnostics.

Empirical StudyLLM AgentsMicroservice Systems

Existing root cause analysis (RCA) research is constrained by Top@k metrics, making it difficult to distinguish between failures originating in the retrieval and reranking stages. This work proposes a retrieval-reranking decoupled framework that constructs a two-stage pipeline comprising a multi-signal fusion retriever and a large language model-based reranker. The proposed approach achieves high-precision root cause localization without requiring causal graphs or annotated data. Experimental results demonstrate that this method comprehensively outperforms the strongest baselines across six benchmarks, improving Top@1 accuracy by up to 18 percentage points.

Anomaly DetectionComplex Monitored SystemsEvaluation Metric

While microservice failures are readily detectable, root cause analysis remains inefficient due to alarm flooding and the absence of structured memory capturing system dependencies and historical behaviors. This work proposes a topology-aware, operation-memory-driven multi-agent architecture that decouples root cause inference from explanation for the first time: the former relies on deterministic computation using a learned dependency graph and temporal anomaly thresholds, while the latter leverages a large language model to generate interpretable recommendations grounded in structured evidence. A novel four-layer operational memory mechanism enables traceable and reusable autonomous operations. Evaluated on an e-commerce benchmark platform with eight types of injected faults, the approach successfully reproduces and resolves two real-world cascading failures, significantly improving diagnostic accuracy and efficiency.

microservice failuresobservabilityoperational memory

Hot Scholars

FD

Franck Dernoncourt

NLP/ML Researcher. MIT PhD.
Machine LearningNeural NetworksNatural Language Processing
IK

Irwin King

The Chinese University of Hong Kong
social computingmachine learningAIgraph neural networks
JC

Jiachi Chen

Associate Professor, Sun Yat-Sen University
Smart ContractsBlockchainLarge Language ModelsSoftware Security
ZZ

Zibin Zheng

IEEE Fellow, Highly Cited Researcher, Sun Yat-sen University, China
BlockchainSmart ContractServices ComputingSoftware Reliability
DT

Dacheng Tao

Nanyang Technological University
artificial intelligencemachine learningcomputer visionimage processing