Understanding Online Failure Prediction in Linux Through Complementary Multi-View Explainability

📅 2026-08-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limited trust in existing online fault prediction methods due to their poor interpretability and susceptibility to workload-specific noise. To overcome these challenges, the authors propose a multi-view complementary and interpretable framework tailored for Linux systems, integrating consensus-based feature selection, temporal onset analysis, subsystem-level causal inference, and multiple diagnostic mechanisms. The framework enables cross-workload evaluation under frozen model conditions. Experimental results demonstrate that the system achieves a detection rate of 91–94% on unseen workloads with a false positive rate below 1%. Furthermore, the study reveals key insights: detection is more robust than diagnosis, warning lead time is highly dependent on fault patterns, and unseen fault patterns cannot be accurately diagnosed solely based on related ones—highlighting the strong sensitivity of diagnosis to specific workloads.
📝 Abstract
Accurate Online Failure Prediction (OFP) has been shown to be feasible in Operating Systems (OSs) settings, but prediction alone is not sufficient for practical adoption. Without diagnostic insight, operators have limited basis to trust alerts or decide how to respond. Moreover, even when predictive accuracy is high, it is often unclear whether models are capturing meaningful failure processes or merely exploiting workload-specific noise and incidental correlations in telemetry. This paper reports a practical experience building and evaluating an explainable OFP pipeline for Linux OSs. We combine consensus-based feature selection for detection with temporal onset analysis, subsystemlevel causal analysis, and complementary diagnostic mechanisms to support failure interpretation. Evaluated under strict crossworkload conditions with frozen training artifacts, it achieved 91-94% detection on unseen workloads without retraining, while maintaining false alarm rates below 1%. However, failure mode diagnosis proved substantially more sensitive to workload shift, and several diagnostics mechanisms showed limited effectiveness for specific failure types. Our experience highlights three main lessons: i) detection generalizes more robustly than diagnosis across workload changes; ii) early-warning capability depends strongly on the failure mode, ranging from 38 to 215 seconds in our study; and iii) unseen failure modes are not reliably diagnosable from related training modes alone, providing 0% accuracy under Leave-One-Mode-Out (LOMO) evaluation. Taken together, these results show the value of complementary explainability mechanisms for interpreting accurate failure predictions, revealing when predictive signals reflect transferable failure structure and when diagnostic generalization breaks down under workload variation.
Problem

Research questions and friction points this paper is trying to address.

Online Failure Prediction
Explainability
Workload Generalization
Failure Diagnosis
Operating Systems
Innovation

Methods, ideas, or system contributions that make the work stand out.

Online Failure Prediction
Explainable AI
Multi-View Explainability
Cross-Workload Generalization
Failure Diagnosis
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
D
Diogo Dória
University of Coimbra, CISUC/LASI, Department of Informatics Engineering, Coimbra, Portugal
J
João R. Campos
University of Coimbra, CISUC/LASI, Department of Informatics Engineering, Coimbra, Portugal