🤖 AI Summary
This study addresses the challenges of workload interference, heterogeneous failure precursors, and non-actionable alerts in fault prediction for industrial GPU clusters by proposing the Falcon framework. Falcon bridges the gap between window-based prediction and practical operational responses through missing-aware temporal feature engineering and peer-relative feature fusion, combined with a fault-specific classifier selection strategy and a threshold-persistent cooldown mechanism. Experimental results demonstrate that Falcon achieves the highest F1 score on the test set, reaching 70.6% for the best-performing fault category, while providing an average early warning lead time of 17 to 35 hours. Furthermore, the proposed framework has been successfully deployed in a production environment, validating its effectiveness for real-world cluster maintenance.
📝 Abstract
GPU clusters are critical infrastructure for AI services, but accurate and actionable GPU failure prediction remains a problem in production settings. We study ticket-linked telemetry from a ByteDance GPU cluster and identify three obstacles: workload-confounded telemetry, heterogeneous fault precursors, and the gap between window-level predictions and actionable alerts. These findings motivate Falcon, a fault-specific warning framework combining missingness-aware temporal and peer-relative features, fault-specific learner selection, and an event policy based on thresholding, persistence, and cooldown. On the test set, Falcon achieves the highest F1 among four baselines and reaches 70.6% F1 on the best-performing fault type. Detected cases provide median lead times of 17.34-35.57 hours. We further report a production deployment, where Falcon is calibrated toward high-precision alerts to reflect false-positive costs. Together, these results show that fault-specific modeling improves early warning from noisy production GPU telemetry.