WavePP: High-Throughput Pipeline Parallel LLM Prefill under Prefix Reuse
This study addresses the high coordination overhead of cross-stage prefix reuse and the limited prefill throughput in pipeline parallelism. To mitigate these issues, this work proposes an asynchronous prefix reuse coordination mechanism coupled with a dynamic chunking strategy. By constructing an efficient runtime that overlaps request admission with execution, the system asynchronously synchronizes cache states across all stages and dynamically plans chunk sizes to maximize pipeline utilization, thereby eliminating admission latency bottlenecks. Implemented on TensorRT-LLM, the proposed approach achieves up to a 2.91× throughput improvement on GLM and MiniMax models, outperforming mainstream baseline systems in most scenarios.