WavePP: High-Throughput Pipeline Parallel LLM Prefill under Prefix Reuse

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high coordination overhead of cross-stage prefix reuse and the limited prefill throughput in pipeline parallelism. To mitigate these issues, this work proposes an asynchronous prefix reuse coordination mechanism coupled with a dynamic chunking strategy. By constructing an efficient runtime that overlaps request admission with execution, the system asynchronously synchronizes cache states across all stages and dynamically plans chunk sizes to maximize pipeline utilization, thereby eliminating admission latency bottlenecks. Implemented on TensorRT-LLM, the proposed approach achieves up to a 2.91× throughput improvement on GLM and MiniMax models, outperforming mainstream baseline systems in most scenarios.
📝 Abstract
Pipeline parallelism can improve prefill throughput by processing multiple request chunks concurrently across different stages of the model. However, keeping the pipeline fully utilized requires efficient scheduling and request preparation. In systems where stages retain and evict cache state independently, a local cache hit does not guarantee that the same prefix can be reused across the pipeline. Here, coordination overhead can impede request admission cadence and thus reduce overall throughput. In this paper, we present WavePP, a prefill runtime built on top of TensorRT-LLM that addresses these challenges by overlapping request admission with pipeline execution. WavePP asynchronously finds a prefix that can be reused across all stages, protects the cached state, and reserves space for the remaining input while earlier requests continue to execute. It subsequently plans the chunk sizes of each request dynamically to maximize pipeline fill. Each stage then completes the local preparation before executing the request. In the same system and pipeline topology, WavePP improves TensorRT-LLM's prefill throughput in 37 of 40 tested settings on GLM 5.2 and MiniMax M2.7. At concurrency 128 with high cache reuse, these changes increase throughput by factors of 2.91 and 2.02, respectively. Across 28 Kimi K3 settings, WavePP also has the highest measured throughput in all 18 settings at concurrency eight or higher, compared with tensor/expert-parallel and pipeline-parallel baselines from TRT-LLM, SGLang, and vLLM.
Problem

Research questions and friction points this paper is trying to address.

Pipeline Parallelism
LLM Prefill
Prefix Reuse
Throughput
Cache Coordination
Innovation

Methods, ideas, or system contributions that make the work stand out.

Pipeline Parallelism
Prefix Reuse
Prefill Throughput
Asynchronous Scheduling
Dynamic Chunking
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Aaryam Sharma
Baseten