Where Root Cause Analysis Fails: A Retrieval-Reranking Decomposition

πŸ“… 2026-09-29
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Existing root cause analysis (RCA) research is constrained by Top@k metrics, making it difficult to distinguish between failures originating in the retrieval and reranking stages. This work proposes a retrieval-reranking decoupled framework that constructs a two-stage pipeline comprising a multi-signal fusion retriever and a large language model-based reranker. The proposed approach achieves high-precision root cause localization without requiring causal graphs or annotated data. Experimental results demonstrate that this method comprehensively outperforms the strongest baselines across six benchmarks, improving Top@1 accuracy by up to 18 percentage points.
πŸ“ Abstract
Identifying the root cause of an anomaly among hundreds of sensors is critical for preventing safety incidents and costly downtime in complex monitored systems. Existing studies evaluate root cause analysis (RCA) methods using top@k accuracy. We show that this metric has a fundamental blind spot: it conflates two failure modes, retrieval failure, where the true cause is never considered, and reranking failure, where it is considered but ranked too low. In this work, we introduce a retrieval-reranking decomposition and audit four well-known benchmarks to expose this blind spot. Our experiments show that, on benchmarks with complex faults, statistical baselines mis-rank the true cause 79-100% of the time, and graph-based methods never clearly beat the best statistical baseline, whether their causal graphs are learned on short fault windows, on retrieved candidate pools guaranteed to contain the cause, or on multi-day normal-operation data. Meanwhile, on simple benchmarks where faults manifest significantly at their origin, retrieval is nearly solved (98-100%). Guided by the decomposition, we build a two-stage pipeline combining a multi-signal retriever with an LLM reranker that, as one fixed configuration, matches or exceeds the best baseline's top@1 accuracy on all six benchmark suites (by up to +12 points), with no causal graph or labeled data required. When all methods rank the same retrieved candidates with the true cause guaranteed present, adding a short system-description document lets the reranker lead the best baseline by +7 to +18 points on every benchmark. Code is available at https://github.com/cruiseresearchgroup/DecompRCA.
Problem

Research questions and friction points this paper is trying to address.

Root Cause Analysis
Anomaly Detection
Retrieval-Reranking Decomposition
Evaluation Metric
Complex Monitored Systems
Innovation

Methods, ideas, or system contributions that make the work stand out.

Root Cause Analysis
Retrieval-Reranking Decomposition
Large Language Model
Two-stage Pipeline
Anomaly Detection
πŸ”Ž Similar Papers
2024-05-03Annual International ACM SIGIR Conference on Research and Development in Information RetrievalCitations: 2
H
Hada Melino Muhammad
University of New South Wales
Luan Pham
Luan Pham
Final-year PhD Candidate @ RMIT, Australia
Software EngineeringAIOpsAnomaly DetectionRoot Cause AnalysisImage Processing
L
Laure Barrière
Baker Hughes
Sachin Shetty
Sachin Shetty
Old Dominion University
BlockchainCyber ResilienceTrustworthy AI
L
Leonardo Pulga
Baker Hughes
F
Flora D. Salim
University of New South Wales