Characterizing Parallelism Strategies in LLM Inference: Fundamental Compute-Communication Trade-offs

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the complexity of parallel strategy selection and the lack of theoretical compute-communication trade-off analysis in large language model (LLM) inference by constructing a unified latency analysis framework. Leveraging distributed system modeling and performance profiling techniques, we develop an end-to-end latency model encompassing computation, communication, and pipeline bubbles. This model precisely characterizes the execution discrepancies among tensor, pipeline, and hybrid parallelism during the prefill and decoding phases, revealing the underlying mechanisms governing the advantages of different strategies. Experimental evaluations validate the accuracy of the proposed model. Ultimately, this work provides robust theoretical guidance for parallel strategy selection and capacity planning in LLM serving systems.
📝 Abstract
Large Language Model (LLM) inference has become the dominant workload in modern AI systems, requiring serving infrastructures to maximize throughput while meeting strict latency Service-Level Objectives (SLOs). Since state-of-the-art LLMs exceed the compute and memory capacity of a single GPU, inference is commonly distributed across multiple GPUs using tensor parallelism (TP), pipeline parallelism (PP), or hybrid parallelism (HB). However, selecting the most effective parallelism strategy remains challenging due to complex interactions among computation, communication, pipeline utilization, sequence length, batch size, and model architecture. Existing approaches largely rely on empirical evaluation and provide limited analytical insight into the trade-offs among these strategies, particularly across the distinct prefill and decoding phases of inference. In this paper, we present a unified analytical framework for modeling distributed LLM inference under TP, PP, and HB. The framework decomposes end-to-end latency into computation, inter-GPU communication, and pipeline bubble overhead, and derives analytical models that capture TP collective communication, PP point-to-point communication, and pipeline utilization as functions of hardware, model, and workload characteristics. The model further characterizes the differing execution behavior of prefill and decoding, explaining why PP-oriented configurations favor compute-intensive prefill while TP-oriented configurations reduce decoding latency by eliminating pipeline bubbles. Experiments with modern LLMs on multi-GPU platforms validate the model and confirm the fundamental compute-communication trade-off across parallelism strategies. The framework provides practical guidance for parallelism selection, capacity planning, and optimization of future LLM serving systems.
Problem

Research questions and friction points this paper is trying to address.

LLM inference
parallelism strategies
compute-communication trade-offs
distributed serving
latency modeling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Language Model Inference
Parallelism Strategies
Analytical Framework
Compute-Communication Trade-off
Pipeline Bubble
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Javad Mirzaei
Office of CTO, Dell Technologies Inc., Canada
Jeebak Mitra
Jeebak Mitra
Dell Technologies
WirelessHigh Speed OpticalDSPAI/MLASIC Design