FluidPD: In-Place Elasticity for SLO-Aware Prefill-Decode Disaggregated LLM Serving

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prefill-decode resource imbalance and SLO violations caused by workload fluctuations in LLM serving. We propose a disaggregated serving system that supports in-place elasticity. The core innovations include FluidToken, a mechanism for temporary token offloading, and FluidRole, which enables in-place role reallocation. By integrating lightweight pressure monitoring with cross-worker computation scheduling, the system achieves adaptive dynamic resource balancing without requiring engine restarts, efficiently handling both transient and sustained load variations. Evaluated on Azure production traces, our approach improves SLO attainment by up to 94.6 percentage points compared to the static SGLang baseline.
📝 Abstract
Prefill-decode disaggregation is becoming a common architecture for LLM serving because it separates two phases with distinct execution patterns and SLO objectives. Existing systems typically combine a fixed prefill/decode worker ratio with request routing across workers. However, real-world workloads exhibit both short bursts and sustained shifts in the prefill-to-decode demand ratio. As a result, a configuration that is well provisioned at one time may quickly become mismatched, causing latency SLO violations even when idle capacity exists elsewhere. Existing autoscaling mechanisms can add capacity, but they react slowly, require spare GPUs, and do not directly address short-timescale phase imbalance. We present FluidPD, a P/D-disaggregated serving system that provides SLO-aware in-place elasticity. FluidPD introduces two complementary mechanisms. FluidToken handles transient imbalance by offloading a bounded portion of prefill computation to decode workers when decode-side slack is available. FluidRole handles sustained imbalance by reassigning running workers between prefill and decode roles in place, avoiding model reload and engine restart. Both mechanisms are guided by lightweight pressure indices that expose prefill and decode-side resource pressure before they appear as SLO violations. Across production Azure trace workloads, FluidPD improves overall SLO attainment over static SGLang by up to 94.6 percentage points, demonstrating that SLO-aware in-place P/D elasticity improves service quality without provisioning additional workers.
Problem

Research questions and friction points this paper is trying to address.

LLM serving
prefill-decode disaggregation
SLO violations
workload imbalance
elasticity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Prefill-Decode Disaggregation
In-Place Elasticity
SLO-Aware Serving
LLM Serving
Autoscaling
🔎 Similar Papers