Longitudinal Analysis of GPU Workloads on Perlmutter

πŸ“… 2025-02-25
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the inefficiency of GPU workloads on the Perlmutter supercomputing platform by conducting the first systematic spatiotemporal analysis based on hardware performance counters. Methodologically, it leverages LDMS to collect fine-grained GPU core and memory performance metrics, and introduces two novel metricsβ€”*burstiness* and *temporal imbalance*β€”to quantify temporal utilization patterns; these are combined with spatial imbalance measures to comparatively analyze usage characteristics of ML and traditional HPC workloads. Results reveal significant spatiotemporal GPU resource imbalance in both workload types: substantial inter-GPU load variance (spatial) and highly volatile utilization over time (temporal), leading to considerable resource underutilization. The findings provide empirical evidence and a quantifiable evaluation framework to guide HPC scheduler optimization, heterogeneous architecture design, and ML–HPC convergence strategies.

Technology Category

Machine Learning: Hardware-aware MLData Mining & Knowledge Management: Mining of Spatial, Temporal or Spatio-Temporal DataPlanning, Routing, and Scheduling: Optimization of Spatio-temporal Systems

Application Category

Systems and Infrastructure for Web, Mobile and WoT: Web performance, measurement, and characterizationGraph Algorithms and Modeling for the Web: Algorithms and analysis for heterogeneous, signed, attributed, multi-relational, temporal, higher-order, and annotated Web-related graphsEconomics, Online Markets and Human Computation: Architectures and workflows that use LLMs for crowd work
πŸ“ Abstract
GPGPU-based clusters and supercomputers have become extremely popular in the last ten years. There is a large number of GPGPU hardware counters exposed to the users, however, very little analysis has been done regarding insights they might offer about workloads running on them. In this work, we address this gap by analyzing previously unexplored GPU hardware counters collected via Lightweight Distributed Metric Service on Perlmutter, a leadership-class supercomputer. We examine several hardware counters related to utilization of GPU cores and memory and present a detailed spatial and temporal analysis of GPU workloads. We investigate spatial imbalance -- uneven GPU usage across multiple GPUs within a job. Our temporal study examines how GPU usage fluctuates during a job's lifetime, introducing two new metrics -- burstiness (the irregularity of large utilization changes) and temporal imbalance (deviations from mean utilization over time). Additionally, we compare machine learning and traditional high performance computing jobs. Our findings uncover inefficiencies and imbalances that can inform workload optimization and future HPC system design.
Problem

Research questions and friction points this paper is trying to address.

Analyzes GPU hardware counters on Perlmutter
Investigates spatial and temporal GPU workload imbalances
Compares machine learning and HPC job efficiencies
Innovation

Methods, ideas, or system contributions that make the work stand out.

Analyzes GPU hardware counters
Introduces burstiness and temporal imbalance
Compares machine learning and HPC jobs