Score
Designs, builds, and operates measurement setups and procedures to quantify power and energy consumed during program or system runtime, including selecting and integrating hardware power meters, sensors, and framework-level probes and instrumenting them for data collection. Executes controlled and normalized workloads, records per-run and per-task energy and power metrics, calibrates and validates instrumentation, and analyzes the resulting consumption data to produce hardware-level energy profiles and reproducible comparisons.
The escalating energy consumption and carbon emissions of software and AI systems necessitate rigorous measurement methodologies. Method: This paper systematically surveys and evaluates existing energy and carbon measurement approaches, proposing the first unified taxonomy classifying methods into monitoring-, estimation-, and black-box-based categories. It conducts a multidimensional assessment across hardware components (CPU, GPU, RAM) and dual dimensions—energy consumption and carbon emissions—grounded in bibliometric analysis and functional comparison of 87 tools. Contribution/Results: Key gaps are identified, including inadequate GPU dynamic power modeling and insufficient carbon intensity mapping for cloud environments. Three pervasive challenges are revealed: poor reproducibility, high hardware heterogeneity, and ambiguous system boundary definitions across the software lifecycle. The findings provide theoretical foundations and practical pathways for establishing standardized benchmarks and advancing green software engineering.
Existing software energy measurement tools struggle to balance accuracy and overhead while often being constrained to specific hardware or programming languages, limiting their cross-platform portability. This work proposes CodeGreen, a modular energy measurement platform that innovatively integrates Tree-sitter–based AST queries to enable automatic, multi-language instrumentation. By decoupling instrumentation from measurement through an asynchronous producer-consumer architecture, CodeGreen supports fine-grained energy analysis for languages including Python, C/C++, and Java. Its Native Energy Measurement Backend (NEMB) unifies polling of hardware sensors such as Intel RAPL, NVIDIA NVML, and AMD ROCm. Evaluated on the Computer Language Benchmarks Game, CodeGreen achieves an energy estimation accuracy with a coefficient of determination of R² = 0.9934 and demonstrates near-perfect workload linearity (R² = 0.9997), offering both high precision and low overhead.
To address the challenges of low accuracy, high overhead, and slow response in online power estimation under enhanced hardware heterogeneity and increased parallelism for embedded systems, this paper proposes a lightweight system-level power modeling and real-time monitoring method based on Performance Monitoring Counters (PMCs). The method constructs a modular, linear-correlation-driven power model that requires no microarchitectural details and supports flexible, rapid reconfiguration across DVFS states. Integrated with the Linux kernel-level framework Runmeter, it enables low-overhead PMC sampling and runtime power estimation. Experimental results demonstrate an average power estimation error of only 7.5%, energy error of 1.3%, and worst-case kernel monitoring overhead below 0.7%. This enables effective closed-loop task scheduling and workload-aware DVFS control.
This work addresses the lack of cost-effective, high-precision power measurement solutions for embedded systems, given the high expense and inflexibility of industrial semiconductor test equipment. The authors propose and implement a compact, open-source hardware and software-based system-level power profiling platform that integrates a Raspberry Pi controller, a high-accuracy current sensor, and a microcontroller-based device under test (DUT). A lightweight HTTP interface enables automated firmware deployment, synchronized execution, and remote control. By uniquely combining low-cost open-source hardware with an automated testing workflow, the platform achieves high-resolution current acquisition and supports energy-efficiency benchmarking and regression testing across multiple firmware variants. This significantly enhances the scalability, reproducibility, and practicality of power analysis for embedded systems, making it well-suited for research, prototyping, and educational applications.
Real-time, low-overhead energy monitoring remains challenging for RISC-V soft-core processors during design space exploration, particularly due to reliance on complex microarchitectural models. Method: This paper proposes a hardware-assisted real-time energy monitoring approach that bypasses such models. It integrates an FPGA system-level module with a custom current/voltage measurement board to directly capture runtime electrical signals, exposing them via a memory-mapped interface for lightweight readout by a monitoring service—achieving high accuracy and low latency without consuming FPGA logic resources. Contribution/Results: The solution supports scalable deployment from single-node to multi-node clusters, enabling synchronized multi-point sampling and distributed analysis. Experimental evaluation demonstrates fine-grained energy-efficiency tracking during RISC-V soft-core execution of shallow neural networks. This provides empirical support for joint performance–energy optimization in power-constrained domains such as aerospace systems.
Processor thermal design power (TDP) is widely misused as a proxy for actual power consumption in physics simulations, leading to inaccurate energy-efficiency assessments. Method: This study conducts the first empirical power and energy measurements of major production-scale physics simulation codes on heterogeneous exascale supercomputers at LLNL and Sandia. Leveraging multi-granularity energy modeling, cross-platform benchmarking, and real-time monitoring across commercial and advanced CPU–GPU heterogeneous nodes, it systematically quantifies runtime energy efficiency. Contribution/Results: Under typical simulation workloads, measured power draw is only 30–60% of TDP—substantially lower than nominal ratings. This work challenges the longstanding practice of substituting TDP for measured power, establishing an empirically grounded methodology for evaluating energy efficiency in exascale systems. It provides critical, reproducible, and generalizable energy benchmarks to guide hardware deployment and energy-aware optimization, thereby advancing low-carbon scientific computing.
This study addresses the lack of a systematic overview of open-source software energy measurement tools, which hinders energy-aware software design and tool selection. From a mining software repositories (MSR) perspective, the authors employ qualitative content analysis to screen and categorize 585 GitHub projects, identifying 24 high-quality open-source energy measurement tools. The work systematically characterizes these tools in terms of architectural design, measurement granularity—spanning from CPU-level to process, container, and AI workload levels—and their capabilities for carbon emission estimation. By elucidating evolutionary trends in tool development, this research provides software architects with a structured foundation and practical guidance for informed tool selection in energy-efficient software engineering.
This study addresses the significant time and energy overhead (0.25%–46.75%) incurred by existing tools for high-frequency RAPL-based power monitoring, which rely on system calls and frequent polling. Through two controlled experiments evaluating seven tools at a 1 kHz sampling rate, the authors develop lightweight user-space applications and kernel modules that directly access model-specific registers (MSRs) using low-level instructions such as rdmsr, bypassing the high-overhead /proc interface. Their findings reveal that system calls are substantially slower than rdmsr, which in turn is slower than common instructions like cpuid. Based on these insights, the work proposes design principles emphasizing architectural simplification and preferential use of low-level instructions, thereby reducing monitoring overhead to near-baseline levels and enabling efficient high-frequency energy analysis.
This work addresses the challenge of achieving efficient and fine-grained module-level power estimation for CPUs during early design stages, where traditional approaches rely on simulation or post-silicon analysis and thus lack agility. The paper introduces, for the first time, the use of large language models for CPU power modeling, proposing a hierarchical source-code-level surrogate model that directly extracts architectural hierarchy, module interconnections, configuration parameters, and workload context from RTL code. This enables accurate per-module power prediction without requiring simulation during inference. Experimental evaluation on the open-source XiangShan processor family demonstrates that the proposed method delivers highly accurate and efficient module-level power estimates across diverse configurations and workloads, significantly outperforming conventional workflows by substantially accelerating early-stage design evaluation while maintaining high fidelity.
This study addresses the challenge of optimizing server energy efficiency in high-throughput computing environments, where performance and energy consumption are often at odds. Leveraging real-world operational data and targeted experiments, the work systematically investigates how server configurations influence power consumption, performance, and carbon emissions, uncovering key barriers to implementing effective energy-saving measures in practice. Through empirical power monitoring, workload modeling, and carbon footprint assessment, the authors identify critical factors governing energy efficiency and propose a practical configuration strategy that simultaneously ensures performance guarantees and advances low-carbon objectives. Evaluated under representative high-throughput workloads, the proposed approach achieves substantial reductions in both energy use and carbon emissions.
This study addresses the limitations of existing AI energy consumption assessments, which typically focus on single inference or training runs and fail to capture the real-world energy dynamics of goal-oriented agent systems involving multi-step execution, retries, and recovery. To bridge this gap, the authors propose the A-LEMS framework, introducing two novel metrics: Energy per successful Goal (EpG) and Orchestration Overhead Index (OOI). A-LEMS integrates a cross-layer observation pipeline with a time-bounded attribution model to enable end-to-end, reproducible energy evaluation. Experimental results demonstrate that agent workflows incur an average EpG of 888.1 joules—4.33 times higher than that of linear baselines—while achieving OOI values below 1.0 in tool-augmented tasks, confirming EpG’s sensitivity and effectiveness in reflecting the energy impact of orchestration structures.