Score
Designs, implements, and evaluates algorithms, firmware, and system-level policies that allocate and enforce power and voltage budgets and perform dynamic power management (e.g., DVFS, power gating) and runtime power control for components and subsystems. Builds energy-aware control, tuning, and optimization methods that analyze and optimize power consumption, power-efficiency, and energy–performance tradeoffs across hardware/software boundaries.
To address the challenges of low accuracy, high overhead, and slow response in online power estimation under enhanced hardware heterogeneity and increased parallelism for embedded systems, this paper proposes a lightweight system-level power modeling and real-time monitoring method based on Performance Monitoring Counters (PMCs). The method constructs a modular, linear-correlation-driven power model that requires no microarchitectural details and supports flexible, rapid reconfiguration across DVFS states. Integrated with the Linux kernel-level framework Runmeter, it enables low-overhead PMC sampling and runtime power estimation. Experimental results demonstrate an average power estimation error of only 7.5%, energy error of 1.3%, and worst-case kernel monitoring overhead below 0.7%. This enables effective closed-loop task scheduling and workload-aware DVFS control.
Traditional dynamic voltage and frequency scaling (DVFS) and default power policies prioritize thermal management and performance while neglecting computational energy efficiency, leading to significant energy waste. To address this, we propose a lightweight, kernel- and power-manager–agnostic energy-saving method leveraging the Linux powercap subsystem and Intel RAPL interface to enforce OS-level power capping—imposing direct power constraints without modifying hardware or firmware. Deployable via a single command, our approach is the first to systematically evaluate power capping against DVFS and default policies in real server environments. On a dual-socket Xeon server, SPEC CPU2017 benchmarks and one-month production measurements demonstrate up to 25% improvement in computational energy efficiency with negligible performance degradation. This work establishes a new paradigm: replacing complex frequency-scaling algorithms with simple, effective power constraints. It provides operations and development teams with a plug-and-play, universally applicable, and highly efficient energy optimization methodology.
This work proposes the first application-agnostic CPU power management approach based on offline reinforcement learning, circumventing the challenges of online methods—such as difficulties in environment modeling, system disturbances, and safety risks. By leveraging historical policy data, the controller is trained without requiring prior knowledge of target applications, utilizing readily available system signals including Intel RAPL power measurements, hardware performance counters, and runtime heartbeat indicators. The method achieves generalizable energy efficiency across diverse workloads while maintaining computational reliability. Experimental evaluation on a range of compute- and memory-intensive benchmarks demonstrates substantial energy savings with only modest and acceptable performance overhead, effectively balancing scientific computing fidelity and energy conservation.
This work addresses the challenge in self-powered streaming networks where dynamic power management, while energy-efficient, incurs switching delays that degrade throughput and hinder real-time signal processing. The paper presents the first cycle-based scheduling framework tailored to such networks, formulating a linear program to compute the maximum achievable throughput and a mixed-integer linear program to minimize energy consumption under throughput constraints. To efficiently explore the trade-off between energy and throughput, the authors introduce a novel “Hop and Skip” multi-objective search strategy that rapidly generates a high-quality Pareto frontier. Experimental results demonstrate that the proposed approach significantly accelerates design space exploration on both benchmark and random graphs, and in practical case studies, it achieves superior energy-throughput trade-offs compared to always-on or purely self-powered baselines.
This work addresses the limitations of traditional dynamic voltage and frequency scaling (DVFS) and task-core allocation methods, which rely on heuristics or offline profiling, struggle to generalize to unseen workloads, and neglect stall time—leading to suboptimal energy efficiency and thermal management. To overcome these challenges, the paper proposes the first large language model (LLM)-guided, zero-shot multi-agent reinforcement learning framework for runtime scheduling. The approach leverages an LLM to extract 13-dimensional code-level semantic features from OpenMP programs and integrates hierarchical multi-agent action decomposition, regression-based environment modeling, and a Dyna-Q architecture to enable workload-agnostic scheduling without prior profiling. Experiments on Jetson TX2/Orin NX, RubikPi, and Intel Core i7 platforms demonstrate a 7.09× improvement in energy efficiency and a 4× reduction in task completion time compared to the Linux ondemand scheduler, with the first scheduling decision made 8,300× faster than conventional tabular methods.
This study addresses the challenge of dynamic scheduling in batteryless IoT systems, where traditional approaches relying on static thresholds or hardware-specific models struggle under highly variable energy availability and workload dynamics. The work proposes two hardware-agnostic, dynamic scheduling strategies that operate without prior knowledge of energy consumption: a model-free reinforcement learning (RL) agent and an online approximate prediction (AP) method, both treating applications as black boxes. To the best of our knowledge, this is the first demonstration of black-box dynamic scheduling in batteryless IoT that is independent of hardware characteristics, systematically uncovering the trade-offs among task throughput, node survivability, and execution pacing. Experimental results show that AP closely approaches oracle-level performance, RL flexibly balances energy usage and survival rate, and adaptive task-rate control (AsTAR) excels under prolonged energy outages, while devices with large capacitors can still benefit from static strategies.
This work addresses the challenge that static power constraints in high-performance computing systems struggle to adapt to the dynamic nature of mixed workloads. To overcome this limitation, the paper proposes a polytopic linear parameter-varying (LPV) H∞ feedback control approach based on workload phase identification. The method employs memory/compute phases as scheduling variables, integrating gain-scheduled PI control with polytopic LPV modeling and optimizing controller performance via H∞ synthesis. Experimental results demonstrate that, under practical power constraints, the proposed LPV controller significantly improves tracking accuracy of the power budget, enhances robustness, and ensures smoother transient behavior during phase transitions compared to conventional gain-scheduled PI controllers.
This work addresses the challenge of achieving efficient and fine-grained module-level power estimation for CPUs during early design stages, where traditional approaches rely on simulation or post-silicon analysis and thus lack agility. The paper introduces, for the first time, the use of large language models for CPU power modeling, proposing a hierarchical source-code-level surrogate model that directly extracts architectural hierarchy, module interconnections, configuration parameters, and workload context from RTL code. This enables accurate per-module power prediction without requiring simulation during inference. Experimental evaluation on the open-source XiangShan processor family demonstrates that the proposed method delivers highly accurate and efficient module-level power estimates across diverse configurations and workloads, significantly outperforming conventional workflows by substantially accelerating early-stage design evaluation while maintaining high fidelity.
This work addresses the neglect of energy efficiency in existing code generation models and the impracticality of large-scale, reproducible hardware-based energy feedback. To bridge this gap, the authors propose the first simulation-based framework for energy-efficient code generation, featuring Green Tea—a deterministic architectural simulator enabling 3.5 million energy evaluations of C++ code snippets. They train energy-aware models via supervised fine-tuning followed by closed-loop reinforcement learning with a novel algorithm (GRPO), and introduce the CARET metric to jointly assess functional correctness and energy efficiency. Experiments on 143 held-out problems demonstrate a 12.63% CARET improvement over baselines, with generated code outperforming human expert implementations in energy efficiency by 58.4%. The study also reveals the misleading nature of traditional throughput-oriented metrics like IPC for energy ranking. The open-sourced dataset and infrastructure eliminate approximately 263,000 CPU hours of reproduction costs.