Crossflow: Prefill-Decode Elasticity for Agentic LLM Serving

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为了解决LLM服务中预填充和解码阶段需求不匹配的问题,Crossflow通过引入弹性边界机制,在不改变节点角色的情况下提高了服务效率和吞吐量。
📝 Abstract
As serving capacity demand surpasses that of training, serving efficiency becomes increasingly important. Prefill-decode (P/D) disaggregation improves serving efficiency through specialization and isolation of the two phases. These benefits rest on a static partitioning. Phase demand, however, is not static. We observe that in a large LLM fleet the ratio of uncached input to output tokens has peak-to-mean ratios up to 4.7x at minute timescales, and that in a public agentic trace the hourly ratio spans a median 24.5x within a single day, while reassigning a replica takes tens of minutes. Agentic traffic sharpens the mismatch. Sizing each pool at its ninety-fifth percentile leaves up to 17% of cluster capacity unused; sizing below it converts the same imbalance into queueing and unrealized throughput. We present Crossflow, which makes this boundary elastic without changing node roles. Each decode node publishes a short-lived, revocable lease that bounds local-prefill compute, KV capacity, transfer work, and projected output. Across public and internal traces, Crossflow improves token throughput by 16.2-17.4% on geometric mean over static P/D, and by up to 43.4% at high load, while reducing mean TTFT at every evaluated point.
Problem

Research questions and friction points this paper is trying to address.

serving efficiency
prefill-decode disaggregation
dynamic demand
resource utilization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Elasticity
Prefill-Decode Disaggregation
Short-lived Lease
Dynamic Adjustment
Agentic LLM Serving
🔎 Similar Papers
No similar papers found.