🤖 AI Summary
GPU reliability degrades under compute-intensive workloads due to aging, yet existing monitoring approaches lack fine-grained, interpretable metrics for early reliability prediction. Method: This paper proposes a holistic stress quantification method that fuses real-time telemetry with low-overhead hardware performance counters—including throughput, instruction issue rate, and pipeline stall events—systematically integrating multi-source signals for parallel workloads. Unlike conventional single-metric monitoring, it constructs a GPU stress assessment model explicitly targeting key functional units (e.g., SMs, memory subsystem). Results: The method accurately characterizes dynamic stress distribution across GPU units, significantly improving early reliability prediction in aging-sensitive scenarios. The proposed stress metric exhibits strong correlation (>0.89) with measured lifetime degradation trends, providing an interpretable, deployable foundation for GPU reliability modeling and runtime health management.
📝 Abstract
Graphics Processing Units (GPUs) are specialized accelerators in data centers and high-performance computing (HPC) systems, enabling the fast execution of compute-intensive applications, such as Convolutional Neural Networks (CNNs). However, sustained workloads can impose significant stress on GPU components, raising reliability concerns due to potential faults that corrupt the intermediate application computations, leading to incorrect results. Estimating the stress induced by an application is thus crucial to predict reliability (with,special,emphasis,on,aging,effects). In this work, we combine online telemetry parameters and hardware performance counters to assess GPU stress induced by different applications. The experimental results indicate the stress induced by a parallel workload can be estimated by combining telemetry data and Performance Counters that reveal the efficiency in the resource usage of the target workload. For this purpose the selected performance counters focus on measuring the i) throughput, ii) amount of issued instructions and iii) stall events.