cross-layer energy measurement

Designs and implements instrumentation and telemetry pipelines to measure and attribute energy consumption across system layers—hardware, compute, network and orchestration—mapping raw hardware signals to higher-level workflows and defining temporal attribution boundaries. Builds reproducible measurement protocols, monitoring agents (including agentic LLM-based collectors), hybrid energy-monitoring systems, and analysis tools that compute metrics such as orchestration-overhead indices and aggregate energy summaries.

cross-layerenergymeasurement

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.2
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

AI infrastructure confronts multidimensional physical and economic constraints—including power, thermal management, water usage, interconnect bandwidth, memory capacity, and data throughput—while existing metrics (e.g., PUE, TCO) are siloed and fail to capture the coupled trade-offs among energy efficiency, performance, and cost, hindering cross-layer co-optimization. To address this, we propose a unified measurement architecture grounded in a 6×3 cross-layer taxonomy—spanning facility, network, compute, storage, software, and application layers, each annotated with physical, computational, and economic semantics—and introduce the Measurement Propagation Graph (MPG) to enable, for the first time, system-level, three-dimensional relational modeling. Leveraging systematic literature review, meta-analysis, and graph-based modeling, our framework integrates heterogeneous, multi-source metrics. It supports benchmarking, capacity planning, and total cost of ownership analysis, substantially enhancing interpretability of AI cluster efficiency frontiers and enabling rigorous multi-objective optimization.

Enables multi-objective optimization of energy, carbon, and costIntegrates physical, computational, and economic constraints into one frameworkUnifies fragmented metrics across AI infrastructure layers

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the limitations of existing AI energy consumption assessments, which typically focus on single inference or training runs and fail to capture the real-world energy dynamics of goal-oriented agent systems involving multi-step execution, retries, and recovery. To bridge this gap, the authors propose the A-LEMS framework, introducing two novel metrics: Energy per successful Goal (EpG) and Orchestration Overhead Index (OOI). A-LEMS integrates a cross-layer observation pipeline with a time-bounded attribution model to enable end-to-end, reproducible energy evaluation. Experimental results demonstrate that agent workflows incur an average EpG of 888.1 joules—4.33 times higher than that of linear baselines—while achieving OOI values below 1.0 in tool-augmented tasks, confirming EpG’s sensitivity and effectiveness in reflecting the energy impact of orchestration structures.

Agentic AIEnergy AccountingEnergy Benchmarking

This study addresses the critical lack of fine-grained, CPU-level energy observability in edge AI devices, which hinders process-level energy attribution and impedes the advancement of low-carbon AI. Through a systematic evaluation of the ASUS Ascent GX10 platform based on the NVIDIA GB10 SoC, the work reveals that the system only supports instantaneous GPU power monitoring and lacks CPU energy counters and standard power management interfaces such as RAPL, thereby failing to replicate the energy tracking capabilities available on x86 platforms. By combining hardware auditing, reverse engineering of ACPI/SPBM, probing of NVML and SCMI protocols, and calibration with external DC power meters, this research uncovers key blind spots in energy observability across mainstream edge AI hardware—further identifying that MediaTek firmware internally computes per-rail energy consumption but does not expose it. The study proposes hardware requirements for energy-attributable AI and advocates integrating energy observability as a core design metric for AI accelerators.

agentic workloadsCPU energy monitoringedge AI

Data-Driven Power Modeling and Monitoring via Hardware Performance Counter Tracking

Jun 30, 2025
SM
Sergio Mazzola
🏛️ ETH Zürich | Real-Time Systems Laboratory (ReTiS) | Scuola Superiore Sant’Anna | RISE Research Institutes of Sweden | Department of Computer Science | Integrated Systems Laboratory (IIS) | Department of Electrical, Electronic, and Information Engineering (DEI) | University of Bologna

To address the challenges of low accuracy, high overhead, and slow response in online power estimation under enhanced hardware heterogeneity and increased parallelism for embedded systems, this paper proposes a lightweight system-level power modeling and real-time monitoring method based on Performance Monitoring Counters (PMCs). The method constructs a modular, linear-correlation-driven power model that requires no microarchitectural details and supports flexible, rapid reconfiguration across DVFS states. Integrated with the Linux kernel-level framework Runmeter, it enables low-overhead PMC sampling and runtime power estimation. Experimental results demonstrate an average power estimation error of only 7.5%, energy error of 1.3%, and worst-case kernel monitoring overhead below 0.7%. This enables effective closed-loop task scheduling and workload-aware DVFS control.

Accurate online power consumption assessment for heterogeneous hardwareDynamic hardware and software adaptation under power constraintsLow-overhead power modeling without microarchitectural details

This study challenges the conventional assumption that device power consumption is predominantly determined by hardware, instead investigating the influence of user behavior on system-level energy usage. Leveraging Intel telemetry data, the research employs exploratory data analysis and linear regression models to compare power consumption patterns across users in different countries, with a focus on the United States and China. The findings reveal a statistically significant association between user behavior and overall power draw, demonstrating that behavioral factors exert a non-negligible impact on energy consumption. This insight offers a novel perspective for green computing initiatives and provides empirical evidence to inform stakeholders such as Intel in refining energy-efficiency strategies and mitigating environmental impact.

energy efficiencyhardware choicepower consumption

This study addresses the lack of a systematic overview of open-source software energy measurement tools, which hinders energy-aware software design and tool selection. From a mining software repositories (MSR) perspective, the authors employ qualitative content analysis to screen and categorize 585 GitHub projects, identifying 24 high-quality open-source energy measurement tools. The work systematically characterizes these tools in terms of architectural design, measurement granularity—spanning from CPU-level to process, container, and AI workload levels—and their capabilities for carbon emission estimation. By elucidating evolutionary trends in tool development, this research provides software architects with a structured foundation and practical guidance for informed tool selection in energy-efficient software engineering.

energy efficiencyenergy measurement toolsGitHub mining

Latest Papers

What's happening recently
View more

This study addresses the challenges of hardware heterogeneity, system complexity, and the absence of unified evaluation benchmarks in edge-to-cloud deployments of agent applications. To this end, it proposes AgenticOps, a framework that enables automated management across the entire agent lifecycle. The framework establishes an end-to-end pipeline integrating distributed deployment with telemetry collection, and introduces an LLM-as-a-Judge mechanism to facilitate semantic-level automated evaluation and reproducible report generation. Experimental results demonstrate that this approach significantly reduces manual operational overhead while providing a standardized experimental paradigm and an efficient evaluation methodology for agent systems.

Agentic ApplicationsEdge-to-Cloud ContinuumHardware Heterogeneity

This study addresses the challenge of optimizing server energy efficiency in high-throughput computing environments, where performance and energy consumption are often at odds. Leveraging real-world operational data and targeted experiments, the work systematically investigates how server configurations influence power consumption, performance, and carbon emissions, uncovering key barriers to implementing effective energy-saving measures in practice. Through empirical power monitoring, workload modeling, and carbon footprint assessment, the authors identify critical factors governing energy efficiency and propose a practical configuration strategy that simultaneously ensures performance guarantees and advances low-carbon objectives. Evaluated under representative high-throughput workloads, the proposed approach achieves substantial reductions in both energy use and carbon emissions.

carbon emissionsdata centersenergy efficiency

This work addresses the challenge of accurately quantifying the carbon footprint of scientific workflows in shared virtualized environments, where existing tools rely on oversimplified power models and lack precision. We propose the first high-fidelity carbon footprint estimation framework that supports multi-cluster deployments and is extensible across diverse workflow systems, including Nextflow and Apache Airflow. Our approach integrates workflow execution traces, node-level fitted power models, hardware-level energy measurements via Intel RAPL, and time-aligned grid carbon intensity data, while accounting for operational emissions, embodied carbon, and water–land resource consumption. Experimental evaluation across three clusters demonstrates an average energy estimation error of only 10.8%, substantially outperforming current tools such as nf-core co2footprint, and confirms successful cross-platform deployment.

carbon footprintenergy consumptionICT emissions

This study addresses the invisibility of energy consumption in software build pipelines and the limitations of existing estimation-based tools that cannot decompose energy usage across individual build phases. To overcome these challenges, this work proposes an open-source command-line tool enabling hardware-level energy measurement with user-defined phase attribution. By monitoring standard output for fine-grained phase-level energy analysis, the approach transcends measurement constraints inherent to cloud environments. Technically, it integrates Intel RAPL, NVML, and Maven plugins with pattern matching to achieve precise energy tracking. Validation on the Gson project successfully generated comprehensive phase-resolved energy profiles for both cold and warm builds. Ultimately, this research provides a reproducible, fine-grained energy efficiency assessment framework to advance green software engineering practices.

build automationCI pipelinesenergy consumption

Hot Scholars

XZ

Xia Zhou

Associate Professor, Columbia University
Mobile computingwireless networkingmobile healthHCI
QL

Qingwen Liu

Tongji University
Wireless NetworkingAI
AV

Alessio Vecchio

University of Pisa
Distributed systemsPervasive computingBioinformatics
SN

Shin Nishio

University College London / Keio University
Quantum Error CorrectionFault Tolerant Quantum ComputingSystem Software