🤖 AI Summary
This study addresses the complexity of parallel strategy selection and the lack of theoretical compute-communication trade-off analysis in large language model (LLM) inference by constructing a unified latency analysis framework. Leveraging distributed system modeling and performance profiling techniques, we develop an end-to-end latency model encompassing computation, communication, and pipeline bubbles. This model precisely characterizes the execution discrepancies among tensor, pipeline, and hybrid parallelism during the prefill and decoding phases, revealing the underlying mechanisms governing the advantages of different strategies. Experimental evaluations validate the accuracy of the proposed model. Ultimately, this work provides robust theoretical guidance for parallel strategy selection and capacity planning in LLM serving systems.
📝 Abstract
Large Language Model (LLM) inference has become the dominant workload in modern AI systems, requiring serving infrastructures to maximize throughput while meeting strict latency Service-Level Objectives (SLOs). Since state-of-the-art LLMs exceed the compute and memory capacity of a single GPU, inference is commonly distributed across multiple GPUs using tensor parallelism (TP), pipeline parallelism (PP), or hybrid parallelism (HB). However, selecting the most effective parallelism strategy remains challenging due to complex interactions among computation, communication, pipeline utilization, sequence length, batch size, and model architecture. Existing approaches largely rely on empirical evaluation and provide limited analytical insight into the trade-offs among these strategies, particularly across the distinct prefill and decoding phases of inference. In this paper, we present a unified analytical framework for modeling distributed LLM inference under TP, PP, and HB. The framework decomposes end-to-end latency into computation, inter-GPU communication, and pipeline bubble overhead, and derives analytical models that capture TP collective communication, PP point-to-point communication, and pipeline utilization as functions of hardware, model, and workload characteristics. The model further characterizes the differing execution behavior of prefill and decoding, explaining why PP-oriented configurations favor compute-intensive prefill while TP-oriented configurations reduce decoding latency by eliminating pipeline bubbles. Experiments with modern LLMs on multi-GPU platforms validate the model and confirm the fundamental compute-communication trade-off across parallelism strategies. The framework provides practical guidance for parallelism selection, capacity planning, and optimization of future LLM serving systems.