Score
Designs and implements models, tools, and measurement methodologies to quantify, estimate, and measure power consumption across abstraction levels (from gate-level to system and data center) and to produce metrics such as power, power/performance, and performance-per-watt. Analyzes and optimizes tradeoffs among power, performance, area, and thermal behavior, and develops sequencing, estimation techniques, and modeling approaches for architectural and system-level power evaluation.
To address the challenges of low accuracy, high overhead, and slow response in online power estimation under enhanced hardware heterogeneity and increased parallelism for embedded systems, this paper proposes a lightweight system-level power modeling and real-time monitoring method based on Performance Monitoring Counters (PMCs). The method constructs a modular, linear-correlation-driven power model that requires no microarchitectural details and supports flexible, rapid reconfiguration across DVFS states. Integrated with the Linux kernel-level framework Runmeter, it enables low-overhead PMC sampling and runtime power estimation. Experimental results demonstrate an average power estimation error of only 7.5%, energy error of 1.3%, and worst-case kernel monitoring overhead below 0.7%. This enables effective closed-loop task scheduling and workload-aware DVFS control.
This work addresses the challenge of achieving efficient and fine-grained module-level power estimation for CPUs during early design stages, where traditional approaches rely on simulation or post-silicon analysis and thus lack agility. The paper introduces, for the first time, the use of large language models for CPU power modeling, proposing a hierarchical source-code-level surrogate model that directly extracts architectural hierarchy, module interconnections, configuration parameters, and workload context from RTL code. This enables accurate per-module power prediction without requiring simulation during inference. Experimental evaluation on the open-source XiangShan processor family demonstrates that the proposed method delivers highly accurate and efficient module-level power estimates across diverse configurations and workloads, significantly outperforming conventional workflows by substantially accelerating early-stage design evaluation while maintaining high fidelity.
Hardware-level power monitoring (e.g., Intel RAPL) suffers from platform dependency and coarse-grained domain-level resolution, hindering fine-grained per-process energy-efficiency analysis. To address this, we propose a hardware-agnostic modeling framework that jointly leverages eBPF and perf to collect fine-grained per-process resource metrics (CPU, memory, I/O, etc.) and integrates node-level power measurements from PDUs. A lightweight regression model is then trained to predict per-process energy consumption with high accuracy. This work presents the first cross-platform, eBPF-driven process–power association model, overcoming the hardware-specificity and granularity limitations of conventional tools. Experimental evaluation demonstrates an average prediction error below 8.3%, significantly enhancing both the precision and interpretability of energy-aware management in data centers.
This study addresses the challenge of optimizing server energy efficiency in high-throughput computing environments, where performance and energy consumption are often at odds. Leveraging real-world operational data and targeted experiments, the work systematically investigates how server configurations influence power consumption, performance, and carbon emissions, uncovering key barriers to implementing effective energy-saving measures in practice. Through empirical power monitoring, workload modeling, and carbon footprint assessment, the authors identify critical factors governing energy efficiency and propose a practical configuration strategy that simultaneously ensures performance guarantees and advances low-carbon objectives. Evaluated under representative high-throughput workloads, the proposed approach achieves substantial reductions in both energy use and carbon emissions.
Processor thermal design power (TDP) is widely misused as a proxy for actual power consumption in physics simulations, leading to inaccurate energy-efficiency assessments. Method: This study conducts the first empirical power and energy measurements of major production-scale physics simulation codes on heterogeneous exascale supercomputers at LLNL and Sandia. Leveraging multi-granularity energy modeling, cross-platform benchmarking, and real-time monitoring across commercial and advanced CPU–GPU heterogeneous nodes, it systematically quantifies runtime energy efficiency. Contribution/Results: Under typical simulation workloads, measured power draw is only 30–60% of TDP—substantially lower than nominal ratings. This work challenges the longstanding practice of substituting TDP for measured power, establishing an empirically grounded methodology for evaluating energy efficiency in exascale systems. It provides critical, reproducible, and generalizable energy benchmarks to guide hardware deployment and energy-aware optimization, thereby advancing low-carbon scientific computing.
This work addresses the challenge in self-powered streaming networks where dynamic power management, while energy-efficient, incurs switching delays that degrade throughput and hinder real-time signal processing. The paper presents the first cycle-based scheduling framework tailored to such networks, formulating a linear program to compute the maximum achievable throughput and a mixed-integer linear program to minimize energy consumption under throughput constraints. To efficiently explore the trade-off between energy and throughput, the authors introduce a novel “Hop and Skip” multi-objective search strategy that rapidly generates a high-quality Pareto frontier. Experimental results demonstrate that the proposed approach significantly accelerates design space exploration on both benchmark and random graphs, and in practical case studies, it achieves superior energy-throughput trade-offs compared to always-on or purely self-powered baselines.
This work addresses the lack of cost-effective, high-precision power measurement solutions for embedded systems, given the high expense and inflexibility of industrial semiconductor test equipment. The authors propose and implement a compact, open-source hardware and software-based system-level power profiling platform that integrates a Raspberry Pi controller, a high-accuracy current sensor, and a microcontroller-based device under test (DUT). A lightweight HTTP interface enables automated firmware deployment, synchronized execution, and remote control. By uniquely combining low-cost open-source hardware with an automated testing workflow, the platform achieves high-resolution current acquisition and supports energy-efficiency benchmarking and regression testing across multiple firmware variants. This significantly enhances the scalability, reproducibility, and practicality of power analysis for embedded systems, making it well-suited for research, prototyping, and educational applications.
Existing GPU power estimation methods often suffer from low accuracy, limited flexibility, or outdated architectural assumptions, making them inadequate for fine-grained energy analysis in modern high-performance computing. To address this gap, this work proposes Wattchmen—a high-fidelity, cross-architecture, instruction-level GPU power modeling framework. By constructing instruction energy models calibrated with diverse microbenchmarks, Wattchmen enables accurate power prediction and attribution across architectures such as V100, A100, and H100 under varying cooling conditions. Experimental evaluation on 16 representative workloads demonstrates that Wattchmen achieves an average absolute percentage error as low as 14% on the V100, substantially outperforming AccelWattch and Guser. Furthermore, it successfully guided energy optimizations for Backprop and QMCPACK, yielding up to 35% energy savings.
This work addresses the critical lack of publicly available, high-resolution power consumption data for generative AI workloads, which hinders accurate data center energy estimation and infrastructure planning. The study presents fine-grained (0.1-second interval) empirical measurements of power draw during training, fine-tuning, and inference tasks on a high-performance computing cluster equipped with NVIDIA H100 GPUs. By integrating standardized benchmarks from MLCommons and vLLM, the authors construct representative workload profiles and extend them to facility-scale energy modeling through an event-driven, bottom-up approach, yielding dynamic power consumption traces that capture real-world temporal fluctuations. This effort delivers the first open, high-resolution dataset on generative AI power usage and introduces a reproducible, scalable methodology to support grid integration and distributed energy resource planning.
Traditional post-layout gate-level power analysis suffers from high computational overhead and poor scalability at sub-clock-cycle granularity, hindering its applicability to efficient power delivery network design and power side-channel security assessments. This work proposes PowerScope—the first machine learning–based framework for sub-cycle power estimation—that directly predicts high-fidelity power waveforms from RTL simulation traces without requiring repeated gate-level simulations. PowerScope establishes the first end-to-end mapping from RTL to sub-cycle power consumption, achieving significant efficiency gains: it attains an average absolute percentage error of 9% (median 5.88%) across diverse benchmarks and operates approximately 80× faster than commercial tools. The framework has been successfully applied to pre-silicon evaluation of power side-channel leakage.