Score
Designs and implements measurement systems, benchmarks, and analyses that quantify power, energy, and carbon impacts of computing workloads and model fine‑tuning across device, node, and system levels. Uses hardware instrumentation and telemetry to collect power and runtime metrics, compute energy-per-event and cost-per-event, perform energy-aware benchmarking and power–performance tradeoff analysis, and produce energy-efficiency characterizations and carbon accounting.
The escalating energy consumption and carbon emissions of software and AI systems necessitate rigorous measurement methodologies. Method: This paper systematically surveys and evaluates existing energy and carbon measurement approaches, proposing the first unified taxonomy classifying methods into monitoring-, estimation-, and black-box-based categories. It conducts a multidimensional assessment across hardware components (CPU, GPU, RAM) and dual dimensions—energy consumption and carbon emissions—grounded in bibliometric analysis and functional comparison of 87 tools. Contribution/Results: Key gaps are identified, including inadequate GPU dynamic power modeling and insufficient carbon intensity mapping for cloud environments. Three pervasive challenges are revealed: poor reproducibility, high hardware heterogeneity, and ambiguous system boundary definitions across the software lifecycle. The findings provide theoretical foundations and practical pathways for establishing standardized benchmarks and advancing green software engineering.
This study addresses the challenge of optimizing server energy efficiency in high-throughput computing environments, where performance and energy consumption are often at odds. Leveraging real-world operational data and targeted experiments, the work systematically investigates how server configurations influence power consumption, performance, and carbon emissions, uncovering key barriers to implementing effective energy-saving measures in practice. Through empirical power monitoring, workload modeling, and carbon footprint assessment, the authors identify critical factors governing energy efficiency and propose a practical configuration strategy that simultaneously ensures performance guarantees and advances low-carbon objectives. Evaluated under representative high-throughput workloads, the proposed approach achieves substantial reductions in both energy use and carbon emissions.
This study addresses the common oversight of full lifecycle carbon emissions in hardware upgrade decisions by proposing a lifecycle-aware simulation framework. The framework uniquely integrates workload characteristics, location-specific time-varying grid carbon intensity, and multiple embodied carbon allocation strategies—such as uniform amortization and front-loading—with multi-generation CPU power models to dynamically evaluate the total carbon footprint of different deployment scenarios. Experimental results demonstrate that, particularly under low-utilization conditions or in regions with cleaner electricity grids, extending the operational lifespan of existing hardware can substantially reduce overall emissions. These findings challenge the prevailing assumption that newer hardware is inherently more environmentally sustainable and offer a novel paradigm for greener computing practices.
To address the challenges of low accuracy, high overhead, and slow response in online power estimation under enhanced hardware heterogeneity and increased parallelism for embedded systems, this paper proposes a lightweight system-level power modeling and real-time monitoring method based on Performance Monitoring Counters (PMCs). The method constructs a modular, linear-correlation-driven power model that requires no microarchitectural details and supports flexible, rapid reconfiguration across DVFS states. Integrated with the Linux kernel-level framework Runmeter, it enables low-overhead PMC sampling and runtime power estimation. Experimental results demonstrate an average power estimation error of only 7.5%, energy error of 1.3%, and worst-case kernel monitoring overhead below 0.7%. This enables effective closed-loop task scheduling and workload-aware DVFS control.
High energy consumption and associated carbon emissions from large-scale scientific workflows in HPC clusters pose a critical sustainability challenge. This paper presents the first cross-disciplinary quantification of carbon footprints across three real-world scientific workflows. We propose an end-to-end energy-efficiency optimization paradigm integrating heterogeneous architecture adaptation, compiler-level energy-aware code optimization, DVFS (Dynamic Voltage and Frequency Scaling), node-level load consolidation, and energy-aware workflow scheduling. Moving beyond single-dimensional optimization, our approach enables holistic hardware–software co-optimization for carbon reduction. Evaluated on representative scientific workflows, it achieves 15–40% reductions in both energy consumption and carbon emissions across the full execution pipeline, significantly improving energy efficiency. The work delivers a reproducible, deployable technical framework and empirically validated benchmarks for green high-performance computing.
Processor thermal design power (TDP) is widely misused as a proxy for actual power consumption in physics simulations, leading to inaccurate energy-efficiency assessments. Method: This study conducts the first empirical power and energy measurements of major production-scale physics simulation codes on heterogeneous exascale supercomputers at LLNL and Sandia. Leveraging multi-granularity energy modeling, cross-platform benchmarking, and real-time monitoring across commercial and advanced CPU–GPU heterogeneous nodes, it systematically quantifies runtime energy efficiency. Contribution/Results: Under typical simulation workloads, measured power draw is only 30–60% of TDP—substantially lower than nominal ratings. This work challenges the longstanding practice of substituting TDP for measured power, establishing an empirically grounded methodology for evaluating energy efficiency in exascale systems. It provides critical, reproducible, and generalizable energy benchmarks to guide hardware deployment and energy-aware optimization, thereby advancing low-carbon scientific computing.
This study addresses the limited scope of traditional high-performance computing (HPC) evaluations, which typically focus solely on performance and energy consumption while overlooking the comprehensive environmental costs of operational configurations. The authors propose the first job-level unified accounting framework that integrates both operational and full life-cycle (embodied) carbon and water footprints. Leveraging life-cycle assessment methodologies, real-time runtime monitoring, and hardware manufacturing emission data, the framework enables fine-grained quantification of environmental impacts. The analysis reveals that increasing thread count generally reduces total environmental footprints, albeit with diminishing marginal returns; while carbon footprints are predominantly driven by operational phases, water footprints are largely dominated by embodied impacts. By jointly incorporating both footprint types at the job granularity, this work establishes a novel paradigm for assessing HPC sustainability.
本文通过引入一个多维度度量框架,解决多尺度高性能计算中的资源管理和可持续性问题,指导现代工作负载的部署策略。
This work addresses the lack of cost-effective, high-precision power measurement solutions for embedded systems, given the high expense and inflexibility of industrial semiconductor test equipment. The authors propose and implement a compact, open-source hardware and software-based system-level power profiling platform that integrates a Raspberry Pi controller, a high-accuracy current sensor, and a microcontroller-based device under test (DUT). A lightweight HTTP interface enables automated firmware deployment, synchronized execution, and remote control. By uniquely combining low-cost open-source hardware with an automated testing workflow, the platform achieves high-resolution current acquisition and supports energy-efficiency benchmarking and regression testing across multiple firmware variants. This significantly enhances the scalability, reproducibility, and practicality of power analysis for embedded systems, making it well-suited for research, prototyping, and educational applications.
High energy consumption has emerged as a critical bottleneck in cloud computing, edge computing, and supercomputing systems supporting large-scale AI applications, necessitating urgent responses to challenges posed by energy use, carbon footprints, and power constraints. This work presents a systematic review of recent advances in energy-efficient computing and introduces the first multidimensional taxonomy that integrates hardware-software co-design, task scheduling, dynamic voltage and frequency scaling, workload consolidation, federated learning, and advanced cooling techniques. Emphasizing the paradigm shift toward green computing in integrated cloud-edge-supercomputing infrastructures driven by large AI models, this study offers a comprehensive roadmap for sustainable computing under carbon constraints, thereby advancing the ICT sector toward greater energy efficiency and lower emissions.
This work addresses the growing impact of carbon emissions from semiconductor device manufacturing and operation as a critical constraint in system design. To tackle this challenge, the paper introduces Architecture Carbon Tool v3 (ACT3), an extensible and customizable sustainability-aware modeling platform that, for the first time, enables carbon-driven design space exploration for silicon-based system architectures. ACT3 significantly enhances modeling capabilities, curated datasets, and analytical telemetry by integrating electronic design automation with architectural modeling techniques, thereby establishing an end-to-end framework for carbon assessment and optimization. Case studies demonstrate that ACT3 effectively supports low-carbon architecture design, offering both a practical tool and novel research directions for sustainable computing systems.