Analytical Power-Aware Provisioning for Prefill-Decode Disaggregated AI Inference

๐Ÿ“… 2026-09-21
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
ๆœฌๆ–‡้’ˆๅฏนAIๆŽจ็†ๆœๅŠกไธญ็š„ๅŠŸ่€—ไธŽๆœๅŠก่ƒฝๅŠ›ๆƒ่กก้—ฎ้ข˜๏ผŒๆๅ‡บไบ†ไธ€็งๅˆ†ๆžๆก†ๆžถ๏ผŒ็”จไบŽๅœจ่€ƒ่™‘็กฌไปถ้™ๅˆถๅ’Œๅทฅไฝœ่ดŸ่ฝฝ็‰นๅพ็š„ๆƒ…ๅ†ตไธ‹ไผ˜ๅŒ–้ข„ๅกซๅ……-่งฃ็ ๅฎžไพ‹็š„ๆ•ฐ้‡ใ€‚
๐Ÿ“ Abstract
Power availability increasingly constrains the operation of AI inference fleets, creating a need for provisioning methods that jointly consider serving capacity and power consumption. Prefill--decode (PD) disaggregation has emerged as a prevalent architecture for large-scale inference serving. However, determining the appropriate numbers of prefill and decode instances is challenging because serving capacity depends jointly on workload characteristics, hardware constraints, queueing, and KV-cache reservations. Existing approaches largely rely on profiling and simulation, providing limited analytical insight into how provisioning decisions shape the tradeoff between serving capacity and power consumption. This paper develops an analytical framework for power-aware provisioning of PD-disaggregated AI inference. Given an inference workload and hardware, the framework models the serving capacity and average power consumption of a provisioned deployment. The serving-capacity model is derived from the joint distribution of input--output lengths and hardware compute and memory limits. In particular, it explicitly captures the coupling between prefill and decode induced by KV-cache reservations, as well as the impact of request queueing. On this basis, the power model determines per-instance power consumption as a function of normalized serving throughput. Together, the models determine the serving capacity--power Pareto front among candidate provisioned deployments, enabling the service provider to choose a provisioned deployment as the workload or available power changes.
Problem

Research questions and friction points this paper is trying to address.

Power-aware provisioning
Prefill-decode disaggregation
Serving capacity
Power consumption
KV-cache
Innovation

Methods, ideas, or system contributions that make the work stand out.

power-aware provisioning
prefill-decode disaggregation
analytical framework
serving capacity
KV-cache
๐Ÿ”Ž Similar Papers
No similar papers found.
M
Mingyuan Yan
Department of Electrical and Computer Engineering, New York University, Brooklyn, New York, USA
H
Haiyu Wang
Department of Electrical and Computer Engineering, New York University, Brooklyn, New York, USA
L
Linxuan Biao
Department of Electrical and Computer Engineering, New York University, Brooklyn, New York, USA
H. Jonathan Chao
H. Jonathan Chao
Professor of ECE, New York University
Networking
Sai Qian Zhang
Sai Qian Zhang
New York University
Wenqi Cui
Wenqi Cui
New York University