🤖 AI Summary
This study investigates how hardware topology, parallelism strategies, and congestion control collectively influence communication exposure and completion efficiency during cross-GPU Mixture-of-Experts (MoE) inference. Leveraging ASTRA-sim integrated with an NS-3 backend and Chakra traces, we construct a deterministic simulation matrix of 768 runs spanning multiple topologies, tensor parallelism (TP) and expert parallelism (EP) partitioning schemes, and network protocols. This work is the first to quantify the joint effects of mapping logical groups onto physical interconnects, revealing that topological advantages are conditional upon the TP degree. Results demonstrate that communication exposure accounts for nearly 95% of completion time. Specifically, TP16EP2 is four times slower than TP2EP16, the DBT algorithm incurs 28%–83% higher latency than Ring, and RoCE DCQCN underperforms InfiniBand HPCC by 23.8%–35.7%.
📝 Abstract
Mixture-of-experts (MoE) models expand capacity via sparse activation, but inference across GPUs introduces tensor-parallel (TP) collectives and expert-parallel (EP) dispatch and combine operations. Completion time depends not just on communication volume but on how logical groups map onto intra-server interconnects, GPU--NIC connections, and the inter-node network. Using ASTRA-sim with the NS-3 discrete-event backend, we build a controlled matrix of 32 GPU ranks with data and pipeline parallelism fixed at one. Workloads are fixed-length 4096-token prefill-like synthetic Chakra traces from four MoE configurations. Experiments cover six server topologies, four TP/EP partitions, two TP collective algorithms, and four network/congestion-control modes, yielding 768 deterministic simulations. In the 144-configuration feedback-enabled subset per model, exposed communication accounts for 89.9%--95.8% of mean completion time. TP16EP2 requires 3.68--4.35x the mean completion time of TP2EP16. With fixed rank mapping, ASTRA-sim Double Binary Tree (DBT) incurs 28.3%--83.2% more time than Ring. InfiniBand-like High Precision Congestion Control (HPCC) is ~0.9% lower than HPCC over RDMA over Converged Ethernet (RoCE), whereas RoCE with Data Center Quantized Congestion Notification (DCQCN) is 23.8%--35.7% slower than RoCE HPCC. Topology effects are conditional: Topology~6 leads at low TP degrees but loses its advantage at high TP degrees, and additional GPUs or NICs help only when rank mapping balances traffic across injection paths. Under uniform 32-way sharding, the largest checkpoint-weight shard is ~48.75 GB per rank, so all configurations meet a 64 GB per-accelerator weight-residency criterion. Within the evaluated workload and simulator semantics, server topology, parallelism, collective implementation, and congestion control jointly determine exposed communication and completion time.