Methodology for Fine-Grain GPU Power Visibility and Insights

📅 2024-12-17
🏛️ arXiv.org
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the challenge of fine-grained power measurement under sub-millisecond AI workloads in GPU clusters for the AI era, this paper proposes FinGraV—a novel methodology for high-fidelity, microsecond-to-millisecond power profiling. FinGraV introduces three core techniques: (1) execution-time binning, (2) hardware-level CPU–GPU precise time synchronization, and (3) differential power profile analysis. Implemented on AMD Instinct MI300X accelerators, it enables accurate, temporally resolved power characterization of key AI operators. For the first time, it systematically reveals dynamic power behaviors across GPU subsystems—including compute units and memory controllers—and uncovers two fundamental phenomena: a strong nonlinear coupling between memory bandwidth and compute-unit power, and a cross-scenario power proportionality law. The methodology is fully reproducible and scalable, establishing a new paradigm for high-performance GPU energy-efficiency modeling and optimization.

Technology Category

Machine Learning: Hardware-aware MLHumans and AI: Other Foundations of Human Computation & AIData Mining & Knowledge Management: Scalability, Parallel & Distributed Systems

Application Category

Security and Privacy: Large-scale security measurementsGraph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphsSystems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applications
📝 Abstract
Ubiquity of AI makes optimizing GPU power a priority as large GPU-based clusters are often employed to train and serve AI models. An important first step in optimizing GPU power consumption is high-fidelity and fine-grain power measurement of key AI computations on GPUs. To this end, we observe that as GPUs get more powerful, the resulting sub-millisecond to millisecond executions make fine-grain power analysis challenging. In this work, we first carefully identify the challenges in obtaining fine-grain GPU power profiles. To address these challenges, we devise FinGraV methodology where we employ execution time binning, careful CPU-GPU time synchronization, and power profile differentiation to collect fine-grain GPU power profiles across prominent AI computations and across a spectrum of scenarios. Using the said FinGraV power profiles, we provide both, guidance on accurate power measurement and, in-depth view of power consumption on state-of-the-art AMD Instinct MI300X. For the former, we highlight a methodology for power differentiation across executions. For the latter, we make several observations pertaining to GPU sub-component power consumption and GPU power proportionality across different scenarios. We believe that FinGraV unlocks both an accurate and a deeper view of power consumption of GPUs and opens up avenues for power optimization of these ubiquitous accelerators.
Problem

Research questions and friction points this paper is trying to address.

Achieving fine-grain GPU power measurement for AI computations
Addressing challenges in sub-millisecond GPU power profiling
Providing insights into GPU sub-component power consumption
Innovation

Methods, ideas, or system contributions that make the work stand out.

Execution time binning for fine-grain power analysis
CPU-GPU time synchronization for accurate measurements
Power profile differentiation across AI computations
🔎 Similar Papers
No similar papers found.
Advanced Micro Devices | MIT
V
Varsha Singhania
Advanced Micro Devices, Inc.
Shaizeen Aga
Shaizeen Aga
AMD Research
Near-data processingSecure hardwareParallel Computer Architecture
M
M. Ibrahim
Advanced Micro Devices, Inc.