GPU Under Pressure: Estimating Application's Stress via Telemetry and Performance Counters

📅 2025-11-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
GPU reliability degrades under compute-intensive workloads due to aging, yet existing monitoring approaches lack fine-grained, interpretable metrics for early reliability prediction. Method: This paper proposes a holistic stress quantification method that fuses real-time telemetry with low-overhead hardware performance counters—including throughput, instruction issue rate, and pipeline stall events—systematically integrating multi-source signals for parallel workloads. Unlike conventional single-metric monitoring, it constructs a GPU stress assessment model explicitly targeting key functional units (e.g., SMs, memory subsystem). Results: The method accurately characterizes dynamic stress distribution across GPU units, significantly improving early reliability prediction in aging-sensitive scenarios. The proposed stress metric exhibits strong correlation (>0.89) with measured lifetime degradation trends, providing an interpretable, deployable foundation for GPU reliability modeling and runtime health management.

Technology Category

Machine Learning: Hardware-aware MLKnowledge Representation and Reasoning: Computational Complexity of ReasoningData Mining & Knowledge Management: Representing, Reasoning, and Using Provenance, Trust

Application Category

Security and Privacy: Large-scale security measurementsSearch and Retrieval-Augmented AI: Web evaluation methodologies and metricsSystems and Infrastructure for Web, Mobile and WoT: Web performance, measurement, and characterization
📝 Abstract
Graphics Processing Units (GPUs) are specialized accelerators in data centers and high-performance computing (HPC) systems, enabling the fast execution of compute-intensive applications, such as Convolutional Neural Networks (CNNs). However, sustained workloads can impose significant stress on GPU components, raising reliability concerns due to potential faults that corrupt the intermediate application computations, leading to incorrect results. Estimating the stress induced by an application is thus crucial to predict reliability (with,special,emphasis,on,aging,effects). In this work, we combine online telemetry parameters and hardware performance counters to assess GPU stress induced by different applications. The experimental results indicate the stress induced by a parallel workload can be estimated by combining telemetry data and Performance Counters that reveal the efficiency in the resource usage of the target workload. For this purpose the selected performance counters focus on measuring the i) throughput, ii) amount of issued instructions and iii) stall events.
Problem

Research questions and friction points this paper is trying to address.

Estimating GPU stress from sustained workloads using telemetry and performance counters
Predicting reliability concerns caused by application-induced component aging
Assessing workload stress through throughput, instructions, and stall metrics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Combining telemetry and performance counters for GPU stress estimation
Using throughput and instruction metrics to assess workload efficiency
Analyzing stall events to predict GPU aging and reliability
G
Giuseppe Esposito
Politecnico di Torino, Dept. of Control and Computer Engineering (DAUIN), Turin, Italy
J
Juan-David Guerrero-Balaguera
Politecnico di Torino, Dept. of Control and Computer Engineering (DAUIN), Turin, Italy
J
J. Condia
Politecnico di Torino, Dept. of Control and Computer Engineering (DAUIN), Turin, Italy
M
M. S. Reorda
Politecnico di Torino, Dept. of Control and Computer Engineering (DAUIN), Turin, Italy
M
Marco Barbiero
AI Delivery Factory, Intesa Sanpaolo S.p.A., Turin, Italy
R
Rossella Fortuna
AI Delivery Factory, Intesa Sanpaolo S.p.A., Turin, Italy