Efficient Expert-Parallel Communication on PCIe-Connected Consumer GPUs

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the communication bottleneck in Mixture-of-Experts (MoE) expert parallelism on consumer-grade PCIe GPUs, where the absence of direct GPU interconnects forces CPU-mediated data transfers, resulting in redundant transmissions and contention for computational resources. To overcome this, we propose ThunderEP, a communication framework tailored for non-directly-connected GPU scenarios. ThunderEP eliminates relay hops inherent in Ring algorithms, leverages DMA engines for data transfer to bypass compute contention, and introduces a low-overhead completion-flag polling mechanism to minimize synchronization latency. The framework has been integrated into the vLLM inference engine. Evaluations on RTX 4090 and RTX 5090 GPUs demonstrate that ThunderEP accelerates the dispatch and combine phases by 2.00× and 1.53×, respectively, achieving up to a 1.66× speedup in end-to-end inference throughput.
📝 Abstract
Expert parallelism (EP) enables inference of large Mixture-of-Experts (MoE) models by placing their experts across multiple GPUs, but requires substantial communication between GPUs at every MoE layer. As contemporary MoE models activate more experts per token, this communication accounts for a growing fraction of inference time. The cost becomes particularly pronounced on PCIe-based consumer GPU systems, where all inter-GPU transfers traverse CPU memory. However, existing MoE-specialized EP communication libraries assume that direct GPU-to-GPU access is available, largely overlooking consumer GPUs. Therefore, most LLM frameworks instead rely on NCCL, whose CPU-staged communication incurs redundant PCIe transfers and competes with expert computation for GPU resources, limiting their overlap. We present ThunderEP, a novel communication design for such systems that removes the relay hops of traditional ring algorithm, moves data through DMA engines to avoid compute resource contention, and minimizes synchronization latency by reducing the polling overhead of completion flags in CPU memory. We integrate the proposed design into vLLM and evaluate it on three widely used MoE models. Experiments on two PCIe systems equipped with RTX 4090 and RTX 5090 GPUs show that ThunderEP achieves average speedups of 2.00$\times$ and 1.53$\times$ over NCCL for dispatch and combine, respectively, and up to 1.66$\times$ end-to-end speedup over state-of-the-art MoE inference frameworks.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
Expert Parallelism
Consumer GPUs
PCIe Communication
LLM Inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

Expert Parallelism
Mixture-of-Experts
Consumer GPUs
DMA Engine
Communication Optimization
🔎 Similar Papers
2024-06-07International Symposium on High-Performance Computer ArchitectureCitations: 5