counterfactual root cause analysis

Designs and implements analytical models and pipelines that identify and quantify causal root causes of observed incidents or faults by generating and evaluating counterfactual incident trajectories and attributing impact to candidate causes. This competence includes estimating causal contributions, ranking and classifying root-cause hypotheses by severity or impact, integrating operator-specified mitigation actions and intervention levels, routing verified hypotheses deterministically, and flagging causes that require experimental validation.

counterfactualrootcauseanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.56
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$214K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Root cause localization in complex, dynamic multi-layer business systems suffers from poor interpretability and difficulty in tracing multi-hop causal dependencies. Method: This paper proposes an end-to-end causal inference framework that uniquely integrates conditional anomaly scoring, counterfactual noise attribution, and depth-first graph search—implemented atop DoWhy—to enable interpretable, backward tracing of multi-hop causal paths from observed anomalies to their initial triggers. Unlike conventional correlation- or rule-based approaches, it reconstructs the full causal chain rather than identifying isolated correlations. Contribution/Results: Evaluated on synthetic anomaly injection benchmarks, the framework achieves significantly higher root cause ranking accuracy than state-of-the-art baselines. It supports actionable root cause diagnosis in dynamic environments by delivering both precise causal attribution and human-interpretable explanations grounded in structural causal models.

Addresses limitations of traditional RCA methods in complex systemsDevelops a causal inference package for actionable root cause analysisIdentifies and ranks root causes by tracing multi-hop causal chains

RCA Copilot: Transforming Network Data into Actionable Insights via Large Language Models

Jul 03, 2025
AS
Alexander Shan
🏛️ Stanford University | Juniper Networks

Traditional root cause analysis (RCA) in complex networked services suffers from high latency, poor interpretability, and heavy reliance on human expertise. To address these limitations, we propose StatLLM-RCA—the first automated RCA framework integrating statistical causal inference with large language models (LLMs). It jointly models multi-source runtime observability data (logs, metrics, traces), applies hypothesis testing to identify plausible causal pathways, and leverages retrieval-augmented generation (RAG) to enable LLMs to produce natural-language attribution reasoning and actionable remediation recommendations. Unlike black-box approaches, StatLLM-RCA ensures fully traceable and verifiable diagnostic reasoning. Experimental evaluation demonstrates that it reduces mean time to identify failures by 58% and achieves a root cause identification accuracy of 92.3%, significantly improving operational decision-making efficiency, transparency, and trustworthiness.

Automating root cause analysis in complex networked servicesEnhancing incident resolution with actionable insights and explanationsOvercoming interpretability challenges in traditional RCA methods

This study addresses the lack of formal definitions in existing root cause analysis methods, which are often limited to root nodes in causal graphs or biased toward proximate causes. Within the potential outcomes framework, this work proposes the first counterfactual definition of root cause at the individual level and introduces a probabilistic measure—Probability of Root Condition (PRC)—to quantify the likelihood that a candidate set of variables constitutes a root cause for a specific outcome. Under standard causal assumptions, the authors derive an explicit identification formula for PRC by integrating causal mediation analysis with counterfactual reasoning, thereby establishing its identifiability. The effectiveness and practical utility of the proposed approach are demonstrated through two numerical examples, filling a critical gap in the formal theory of root cause analysis.

causal inferencecounterfactualpotential outcomes

Causal AI-based Root Cause Identification: Research to Practice at Scale

Feb 25, 2025
SJ
Saurabh Jha
🏛️ IBM | Guild Systems Inc.

Modern distributed systems suffer from frequent failures, and conventional APM tools—relying heavily on correlation analysis—exhibit high false-positive rates and poor interpretability, hindering near-real-time root-cause localization. To address this, we propose a novel root-cause identification paradigm grounded in causal AI. Our approach introduces the first end-to-end inference framework that jointly integrates structural causal models (SCMs), dynamic Bayesian networks, and streaming causal discovery to enable real-time causal reasoning over heterogeneous, multi-source monitoring data. The method has been integrated into IBM Instana and deployed at scale in enterprise production environments. Empirical evaluation demonstrates that it reduces mean root-cause localization time to the sub-second level and improves accuracy by 42%, significantly enhancing system reliability and SLO compliance assurance.

Enhance system reliability using causal AIIdentify root causes in distributed systemsImplement causal RCI algorithm at scale

Root Cause Analysis of Outliers with Missing Structural Knowledge

Jun 07, 2024
NO
Nastaran Okati
🏛️ Max Planck Institute for Software Systems | Max Planck Institute for Intelligent Systems | University of Cambridge | Amazon Research

Real-world root cause analysis (RCA) faces a critical challenge: post-intervention distributions often contain only a few—or even a single—sample, rendering distribution-dependent or low-density-region regression methods statistically ill-posed. This paper proposes a lightweight root cause identification framework that requires neither counterfactual reasoning nor a fully specified structural causal model (SCM). It operates either given a causal DAG or, in the absence of one, solely from an anomaly score ranking. We theoretically prove that low-scoring anomalies rarely trigger high-scoring ones and derive a probabilistic upper bound on non-monotonic propagation paths. By abandoning Shapley-value-based attribution and density-sensitive regression, our method achieves linear time complexity O(n). It eliminates SCM fitting and counterfactual computation while providing rigorous theoretical guarantees and strong empirical performance.

Addresses single-sample limitations in post-intervention distribution analysisIdentifies root causes of anomalies with missing causal graph knowledgeProvides guarantees for root cause detection in polytree structures

Latest Papers

What's happening recently
View more

Industrial root cause diagnosis typically relies on manual hypotheses and extensive fault labels, yet existing data-driven methods suffer from poor interpretability and limited generalization. This work proposes AgentRCA, a novel framework that achieves zero-shot, label-free root cause diagnosis for the first time. By integrating data-driven digital twins with tool-augmented large language models, AgentRCA adopts a hypothesis-driven approach to iteratively gather statistical evidence, evaluate competing hypotheses, and construct transparent reasoning chains that explicitly link observed symptoms to underlying physical faults. Evaluated on real-world multiphase flow facilities and large-scale chemical plants, the framework matches the diagnostic performance of fully supervised baselines while offering high interpretability.

anomaly diagnosisexplainable AIindustrial operation

This work addresses the limited interpretability and accountability of large language models (LLMs) in root cause analysis, which hinder their applicability in high-stakes operational settings requiring rigorous evidence chains, hypothesis comparison, and uncertainty handling. The authors propose JustDiag, a diagnostic argumentation engine that introduces, for the first time, an explicit modeling of the diagnostic reasoning process into root cause analysis. JustDiag structures and maintains states such as evidence, findings, competing hypotheses, conflicts, and follow-up checks to enable traceable and auditable inference, complemented by a calibration mechanism that explicitly accounts for uncertainty. Integrating LLMs with a structured reasoning framework, the approach employs a two-tier evaluation protocol to assess both outcome and reasoning quality. Experiments on 66 real-world incidents demonstrate that JustDiag significantly outperforms non-argumentative baselines in both outcome and process scores, exhibiting superior uncertainty retention despite a slightly lower completion rate.

accountabilitydiagnostic justificationincident response

This work addresses the limitations of existing data-driven root cause analysis methods, which rely on the causal sufficiency assumption and suffer significant performance degradation in partially observable systems with unmeasured latent variables. The authors propose a novel approach based on partial ancestral graphs (PAGs), modeling system failures as parametric interventions and integrating causal effect identification with partial identification theory to rank candidate root causes. Notably, for non-identifiable scenarios, the method introduces analytical causal bounds for the first time in root cause analysis. This framework is the first to jointly handle latent variables and partial identifiability, thereby eliminating dependence on causal sufficiency. Experiments demonstrate that the proposed method substantially outperforms state-of-the-art approaches across synthetic data, microservice anomaly benchmarks, and power grid cascading failure datasets, confirming its robustness and effectiveness in complex, partially observable environments.

Causal SufficiencyLatent ConfoundersPartial Ancestral Graphs

Hot Scholars

GW

Guancheng Wang

the Research Ireland Centre for Software, University of Limerick
Software Testing and Debugging
LC

Liqian Chen

Professor, National University of Defense Technology, China
Program analysis & verificationAbstract interpretationProgram repair
KZ

Keyi Zhang

Lead Compiler Engineer at Efficient Computer
AW

Andreas Wiedholz

Researcher at XITASO GmbH
Roboticsartificial intelligencecomputer vision