VarioPath: Workload-Aware All-to-All Communication for PCIe GPU Clusters

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inefficiency of AlltoAllv communication in PCIe clusters caused by link contention and traffic skew, proposing the VarioPath framework. VarioPath introduces a novel synergistic mechanism combining offline topology analysis with online demand scheduling. Specifically, the offline phase leverages symmetry compression to construct a catalog of contention-free channels, while the online phase achieves dynamic path scheduling through adaptive traffic decomposition, significantly reducing computational overhead. Experimental results demonstrate that VarioPath attains an average 5.88× communication speedup, reduces inference latency by 27.2% for Qwen3, and decreases generation latency by 6.1% for Wan2.1.
📝 Abstract
AlltoAllv communication is a critical primitive in distributed large-model inference, particularly for mixture-of-experts (MoE) models. The growing adoption of PCIe GPU systems for cost-efficient inference makes AlltoAllv performance on these systems increasingly important. Without a dedicated scale-up interconnect (e.g., NVLink or Infinity Fabric), PCIe GPU systems carry both intra-node and inter-node traffic through the PCIe hierarchy, where concurrent transfers can contend for PCIe link bandwidth. This link contention, compounded by skewed traffic distributions and dynamic traffic demand, makes efficient AlltoAllv scheduling challenging. Existing approaches are either poorly suited to PCIe GPU systems or incur substantial schedule synthesis overhead that reduces their practicality in real-world deployments. We present VarioPath, an efficient AlltoAllv scheduling framework for PCIe GPU systems. It combines an offline topology-aware analyzer with an online demand-aware scheduler. The analyzer records contention-free transfer patterns as AlltoAllv channels and exploits topology symmetry to build a compact catalog for efficient search. The online scheduler decomposes each AlltoAllv invocation's demand across a sequence of channels, adapting to rapidly changing and skewed traffic while incurring low planning overhead. Evaluation on four platforms (up to 256 GPUs) shows average AlltoAllv speedups of 5.88x over FAST and 1.72x over DeepEP. End-to-end experiments show that VarioPath reduces Qwen3 inference latency by up to 27.2% and Wan2.1 generation latency by 6.1%.
Problem

Research questions and friction points this paper is trying to address.

AlltoAllv communication
PCIe GPU clusters
link contention
distributed inference
mixture-of-experts
Innovation

Methods, ideas, or system contributions that make the work stand out.

AlltoAllv scheduling
PCIe GPU clusters
workload-aware
topology-aware analyzer
Mixture-of-Experts
🔎 Similar Papers
2024-06-07International Symposium on High-Performance Computer ArchitectureCitations: 5