🤖 AI Summary
This work addresses the high power consumption of modern GPUs in machine learning data centers, a challenge exacerbated by existing data-center-level optimization approaches that overlook fine-grained power variations among internal GPU components. The paper proposes CompPow, a novel methodology that systematically demonstrates the critical role of component-level power awareness in improving energy efficiency. By constructing a fine-grained component power model and integrating it with analysis of ML workload characteristics, CompPow enables hardware-software co-optimized, component-level power management. Evaluated across diverse machine learning operations and execution modes, the approach achieves up to 10% higher energy efficiency and 5% performance improvement, establishing a new paradigm for GPU-centric, co-designed energy-efficient computing.
📝 Abstract
The ever increasing demand for ML-driven intelligence in a wide spectrum of domains has led to ubiquity of GPUs. At the same time, GPUs are notorious for their power consumption needs and often dominate power allocation in a typical ML datacenter. While datacenter-level power optimizations which focus on collection of GPUs are promising, in this work, we take a different tack -- namely, we take a closer look at power consumption inside a GPU. Specifically, as modern GPUs are comprised of integrated components, we make a case for component-awareness, termed CompPow in this work, for improved power management in modern GPUs. We demonstrate for a variety of ML operations and execution patterns, CompPow has the potential to deliver higher energy efficiency (10%) and even improved performance (5%). We conclude with recommendations on how component-aware software-hardware co-design can extract additional energy efficiency from modern GPUs.