From Noisy Telemetry to Actionable Warnings: GPU Failure Prediction in Industrial Clusters

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of workload interference, heterogeneous failure precursors, and non-actionable alerts in fault prediction for industrial GPU clusters by proposing the Falcon framework. Falcon bridges the gap between window-based prediction and practical operational responses through missing-aware temporal feature engineering and peer-relative feature fusion, combined with a fault-specific classifier selection strategy and a threshold-persistent cooldown mechanism. Experimental results demonstrate that Falcon achieves the highest F1 score on the test set, reaching 70.6% for the best-performing fault category, while providing an average early warning lead time of 17 to 35 hours. Furthermore, the proposed framework has been successfully deployed in a production environment, validating its effectiveness for real-world cluster maintenance.
📝 Abstract
GPU clusters are critical infrastructure for AI services, but accurate and actionable GPU failure prediction remains a problem in production settings. We study ticket-linked telemetry from a ByteDance GPU cluster and identify three obstacles: workload-confounded telemetry, heterogeneous fault precursors, and the gap between window-level predictions and actionable alerts. These findings motivate Falcon, a fault-specific warning framework combining missingness-aware temporal and peer-relative features, fault-specific learner selection, and an event policy based on thresholding, persistence, and cooldown. On the test set, Falcon achieves the highest F1 among four baselines and reaches 70.6% F1 on the best-performing fault type. Detected cases provide median lead times of 17.34-35.57 hours. We further report a production deployment, where Falcon is calibrated toward high-precision alerts to reflect false-positive costs. Together, these results show that fault-specific modeling improves early warning from noisy production GPU telemetry.
Problem

Research questions and friction points this paper is trying to address.

GPU failure prediction
noisy telemetry
industrial clusters
actionable warnings
fault precursors
Innovation

Methods, ideas, or system contributions that make the work stand out.

GPU failure prediction
fault-specific modeling
actionable warnings
telemetry data
event policy
🔎 Similar Papers
No similar papers found.