Score
Designs and implements analytical models and pipelines that identify and quantify causal root causes of observed incidents or faults by generating and evaluating counterfactual incident trajectories and attributing impact to candidate causes. This competence includes estimating causal contributions, ranking and classifying root-cause hypotheses by severity or impact, integrating operator-specified mitigation actions and intervention levels, routing verified hypotheses deterministically, and flagging causes that require experimental validation.
Existing root cause analysis (RCA) research lacks a goal-oriented, systematic taxonomy, leading to task ambiguity and hindered progress assessment. Method: This paper proposes the first RCA classification framework centered on fundamental objectives—departing from conventional data-type–based taxonomies—and systematically categorizes 135 studies (2014–2025) according to core goals such as fault localization and defect remediation. Guided by a systematic literature review, we construct a multi-level RCA objective hierarchy that characterizes the state of the art, recurrent challenges, and critical technical gaps per task. Contribution/Results: We present the first RCA objective-method mapping atlas tailored to cloud service scenarios, establishing a theoretical foundation for academic research and a practical technology roadmap for industrial deployment.
Root cause localization in complex, dynamic multi-layer business systems suffers from poor interpretability and difficulty in tracing multi-hop causal dependencies. Method: This paper proposes an end-to-end causal inference framework that uniquely integrates conditional anomaly scoring, counterfactual noise attribution, and depth-first graph search—implemented atop DoWhy—to enable interpretable, backward tracing of multi-hop causal paths from observed anomalies to their initial triggers. Unlike conventional correlation- or rule-based approaches, it reconstructs the full causal chain rather than identifying isolated correlations. Contribution/Results: Evaluated on synthetic anomaly injection benchmarks, the framework achieves significantly higher root cause ranking accuracy than state-of-the-art baselines. It supports actionable root cause diagnosis in dynamic environments by delivering both precise causal attribution and human-interpretable explanations grounded in structural causal models.
Traditional root cause analysis (RCA) in complex networked services suffers from high latency, poor interpretability, and heavy reliance on human expertise. To address these limitations, we propose StatLLM-RCA—the first automated RCA framework integrating statistical causal inference with large language models (LLMs). It jointly models multi-source runtime observability data (logs, metrics, traces), applies hypothesis testing to identify plausible causal pathways, and leverages retrieval-augmented generation (RAG) to enable LLMs to produce natural-language attribution reasoning and actionable remediation recommendations. Unlike black-box approaches, StatLLM-RCA ensures fully traceable and verifiable diagnostic reasoning. Experimental evaluation demonstrates that it reduces mean time to identify failures by 58% and achieves a root cause identification accuracy of 92.3%, significantly improving operational decision-making efficiency, transparency, and trustworthiness.
This study addresses the lack of formal definitions in existing root cause analysis methods, which are often limited to root nodes in causal graphs or biased toward proximate causes. Within the potential outcomes framework, this work proposes the first counterfactual definition of root cause at the individual level and introduces a probabilistic measure—Probability of Root Condition (PRC)—to quantify the likelihood that a candidate set of variables constitutes a root cause for a specific outcome. Under standard causal assumptions, the authors derive an explicit identification formula for PRC by integrating causal mediation analysis with counterfactual reasoning, thereby establishing its identifiability. The effectiveness and practical utility of the proposed approach are demonstrated through two numerical examples, filling a critical gap in the formal theory of root cause analysis.
Modern distributed systems suffer from frequent failures, and conventional APM tools—relying heavily on correlation analysis—exhibit high false-positive rates and poor interpretability, hindering near-real-time root-cause localization. To address this, we propose a novel root-cause identification paradigm grounded in causal AI. Our approach introduces the first end-to-end inference framework that jointly integrates structural causal models (SCMs), dynamic Bayesian networks, and streaming causal discovery to enable real-time causal reasoning over heterogeneous, multi-source monitoring data. The method has been integrated into IBM Instana and deployed at scale in enterprise production environments. Empirical evaluation demonstrates that it reduces mean root-cause localization time to the sub-second level and improves accuracy by 42%, significantly enhancing system reliability and SLO compliance assurance.
Real-world root cause analysis (RCA) faces a critical challenge: post-intervention distributions often contain only a few—or even a single—sample, rendering distribution-dependent or low-density-region regression methods statistically ill-posed. This paper proposes a lightweight root cause identification framework that requires neither counterfactual reasoning nor a fully specified structural causal model (SCM). It operates either given a causal DAG or, in the absence of one, solely from an anomaly score ranking. We theoretically prove that low-scoring anomalies rarely trigger high-scoring ones and derive a probabilistic upper bound on non-monotonic propagation paths. By abandoning Shapley-value-based attribution and density-sensitive regression, our method achieves linear time complexity O(n). It eliminates SCM fitting and counterfactual computation while providing rigorous theoretical guarantees and strong empirical performance.
Industrial root cause diagnosis typically relies on manual hypotheses and extensive fault labels, yet existing data-driven methods suffer from poor interpretability and limited generalization. This work proposes AgentRCA, a novel framework that achieves zero-shot, label-free root cause diagnosis for the first time. By integrating data-driven digital twins with tool-augmented large language models, AgentRCA adopts a hypothesis-driven approach to iteratively gather statistical evidence, evaluate competing hypotheses, and construct transparent reasoning chains that explicitly link observed symptoms to underlying physical faults. Evaluated on real-world multiphase flow facilities and large-scale chemical plants, the framework matches the diagnostic performance of fully supervised baselines while offering high interpretability.
This work addresses the limited interpretability and accountability of large language models (LLMs) in root cause analysis, which hinder their applicability in high-stakes operational settings requiring rigorous evidence chains, hypothesis comparison, and uncertainty handling. The authors propose JustDiag, a diagnostic argumentation engine that introduces, for the first time, an explicit modeling of the diagnostic reasoning process into root cause analysis. JustDiag structures and maintains states such as evidence, findings, competing hypotheses, conflicts, and follow-up checks to enable traceable and auditable inference, complemented by a calibration mechanism that explicitly accounts for uncertainty. Integrating LLMs with a structured reasoning framework, the approach employs a two-tier evaluation protocol to assess both outcome and reasoning quality. Experiments on 66 real-world incidents demonstrate that JustDiag significantly outperforms non-argumentative baselines in both outcome and process scores, exhibiting superior uncertainty retention despite a slightly lower completion rate.
This work addresses the limitations of existing data-driven root cause analysis methods, which rely on the causal sufficiency assumption and suffer significant performance degradation in partially observable systems with unmeasured latent variables. The authors propose a novel approach based on partial ancestral graphs (PAGs), modeling system failures as parametric interventions and integrating causal effect identification with partial identification theory to rank candidate root causes. Notably, for non-identifiable scenarios, the method introduces analytical causal bounds for the first time in root cause analysis. This framework is the first to jointly handle latent variables and partial identifiability, thereby eliminating dependence on causal sufficiency. Experiments demonstrate that the proposed method substantially outperforms state-of-the-art approaches across synthetic data, microservice anomaly benchmarks, and power grid cascading failure datasets, confirming its robustness and effectiveness in complex, partially observable environments.