Do RUL explanations hold up? Faithfulness and stability of attributions on C-MAPSS

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing explanation methods for deep remaining useful life (RUL) prediction models in terms of faithfulness and stability evaluation. For the first time, it systematically evaluates deletion/insertion faithfulness and cross-seed stability, explicitly distinguishing noise robustness from retraining consistency. Based on CNN, LSTM, and Transformer architectures, this work comparatively analyzes the explainable AI (XAI) performance of attribution algorithms, including Integrated Gradients, Occlusion, and attention mechanisms. The results demonstrate that Occlusion and Integrated Gradients achieve superior faithfulness, rendering them suitable for practical engineering maintenance reports, whereas raw attention should serve solely as a visualization aid. These findings provide a reliable foundation for interpretable predictive maintenance.
📝 Abstract
Deep remaining-useful-life (RUL) models on NASA C-MAPSS are now routine, and so are heatmaps that colour sensors and timesteps. A heatmap that looks mechanical is not the same as an explanation an engineer can act on. We train three standard architectures - a 1D CNN, an LSTM, and a small Transformer encoder - on the official FD001 and FD003 splits with the piecewise RUL cap of 125 cycles and the official PHM08 asymmetric score. We then attach three attribution maps (Integrated Gradients, occlusion, last-layer attention) and evaluate them with the checks the XAI-for-PdM literature still under-reports: deletion/insertion faithfulness, Spearman stability under sensor-scale noise, agreement across training seeds, and cosine consistency inside RUL bins. Prediction error is a prerequisite, not the claim. The headline is which explanation method moves the RUL output when its top cells are removed, and which map survives a 5% input perturbation. Integrated Gradients and occlusion are similarly faithful on the LSTM; Transformer attention is cheap and temporally smooth but weakly faithful. All three maps are almost unchanged under 5% input noise, yet IG/occlusion agree only moderately across two LSTM seeds - stability to sensor jitter is not the same as stability to retraining. A secondary tabular check on the AI4I 2020 failure dataset shows the same deletion pattern for tree importances. We recommend occlusion or IG for any C-MAPSS-style report that will be read by a maintenance engineer, and we treat raw attention weights as a visualisation only.
Problem

Research questions and friction points this paper is trying to address.

Remaining Useful Life
Attribution Faithfulness
Explanation Stability
Predictive Maintenance
Explainable AI
Innovation

Methods, ideas, or system contributions that make the work stand out.

Remaining Useful Life (RUL)
Attribution Methods
Faithfulness Evaluation
Stability Analysis
Explainable AI (XAI)
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Manh Hien Nguyen
AI Lab, Phuong Hai JSC, Vietnam
Ngoc Thanh Nguyen
Ngoc Thanh Nguyen
International College of Management Sydney, Australia
I
Isabella Mendoza Cortes
International College of Management Sydney, Australia
T
Tam Khuat
International College of Management Sydney, Australia
T
Thanh Pham
School of Science, Engineering, and Technology, RMIT, Vietnam
N
Nhat Quang Tran
School of Science, Engineering, and Technology, RMIT, Vietnam
U
Ushik Shrestha Khwakhali
School of Science, Engineering, and Technology, RMIT, Vietnam
L
Loan Do
FPT, Vietnam