ePACT: Energy-Performance-Aware Commitment Tracking for LLM Serving

πŸ“… 2026-10-01
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the escalating deployment costs in large language model (LLM) serving caused by power commitment deviations, proposing the ePACT framework. This method introduces an asymmetric deviation cost model and employs a bi-level controller to dynamically coordinate GPU capacity and clock frequency adjustments. By integrating global planning with local decision-making, it achieves coarse-to-fine asynchronous action search to optimize the trade-off between energy consumption and performance. Experiments conducted on vLLM with H20/H200 GPU pools demonstrate that, over 24-hour simulations, ePACT reduces asymmetric deviation costs by 73.8%–75.7% while maintaining service level objective attainment rates comparable to native vLLM.
πŸ“ Abstract
Reducing LLM serving energy does not by itself guarantee lower deployment cost when electricity procurement exposes operators to unfavorable deviations from preset commitments. We study hourly commitments with positive, potentially asymmetric costs for overuse and underuse, and formulate energy-Performance-Aware Commitment Tracking: minimize deviation costs subject to request-level service requirements. We implement ePACT, a two-level controller that adjusts serving capacity and GPU clocks as requests arrive. A global planner updates interval energy targets from measured consumption and the remaining hourly commitment. A local decision maker predicts candidate configurations' energy and completion times, checks predicted deadline misses, and selects among admitted configurations by asymmetric target-deviation cost, with a service-first fallback. Coarse-to-fine action search runs asynchronously with serving. We evaluate ePACT through single-hour comparisons, controller ablations, and full-day trace simulations for H20 and H200 GPU pools. In the 24-hour simulations, ePACT reduces the asymmetric deviation cost by $73.8\%$ and $75.7\%$ relative to vLLM while retaining near-vLLM SLO attainment. Mean absolute hourly deviations are $2.16\%$ and $2.31\%$, respectively.
Problem

Research questions and friction points this paper is trying to address.

LLM serving
energy commitment tracking
deviation cost minimization
service level objectives
asymmetric cost
Innovation

Methods, ideas, or system contributions that make the work stand out.

Commitment Tracking
LLM Serving
Energy-Performance Tradeoff
Two-level Controller
GPU Clock Scaling
πŸ”Ž Similar Papers
2024-08-05International Symposium on High-Performance Computer ArchitectureCitations: 5