🤖 AI Summary
This work addresses the limitations of existing root cause analysis (RCA) methods, which rely on strong assumptions and struggle with latent confounders and complex distributional constraints. We propose the Deep Causal Model (DCM) framework, which represents systems as causal graphs and leverages deep neural networks to parameterize variations in structural functions. By transcending traditional conditional independence restrictions, DCM implicitly connects distributional constraint testing with RCA, enabling the integration of arbitrary causal models and partial graph knowledge for precise fault localization under incomplete observations. Experimental evaluations on the Sock Shop and Online Boutique datasets demonstrate that DCM significantly outperforms baseline methods in Top-1 accuracy. Notably, in the presence of unobserved confounders, the proposed approach achieves a perfect recovery rate of 0.846, substantially surpassing state-of-the-art techniques.
📝 Abstract
Root cause analysis (RCA) is a critical problem in many real-world scenarios. RCA enables the identification of faulty or failing mechanisms in a system by comparing anomalous observations with corresponding reference (i.e., regular) observations. However, existing approaches rely either on heuristic methods or on conditional independence tests with a strong unconfoundedness assumption, and thus fail to exploit other complicated distributional constraints in the presence of latent variables. To relax these assumptions, we model the underlying system as a causal model and the anomalous system as a change in the structural functions of the same causal model. Specifically, to handle unobserved confounders, we establish an implicit connection between distributional constraint testing and root cause analysis. To adapt our approach to data generated from arbitrary causal models, we employ the deep causal model (DCM) framework, in which we design the causal model using neural networks. Finally, we illustrate how our method, RCA-DCM, can utilize different levels of partial graphical knowledge to perform RCA. We evaluate RCA-DCM against state-of-the-art baselines on simulated datasets, a physics-based causal chamber and two micro-service applications. RCA-DCM improves top-1 accuracy over the strongest baseline on both Sock Shop (0.880 vs. 0.752) and Online Boutique (0.776 vs. 0.712), and when the true root cause in the causal chamber is unobserved and acts as a latent confounder, it recovers the exact root-cause set more often than any competing method (perfect recovery rate (PRR) 0.846 vs. 0.731).