🤖 AI Summary
This study addresses the severe communication bottlenecks encountered during Mixture-of-Experts (MoE) inference on consumer-grade GPUs due to the absence of high-bandwidth interconnects. To this end, we propose CoMoE, a system that introduces a novel host-centric routing architecture, leveraging the host as an active routing hub to optimize token dispatch and aggregation. Its core innovations include host-endorsed token multicasting to eliminate redundant transmissions and a scratchpad-buffer-based fine-grained aggregation mechanism that circumvents global synchronization stalls. Experimental results demonstrate that CoMoE improves inference throughput by up to 1.46× on RTX 5090 GPUs, achieving performance comparable to data-center GPUs at merely 23.4% of their hardware cost. This work thereby presents an efficient solution for low-cost MoE deployment.
📝 Abstract
Deploying Mixture-of-Experts (MoE) models relies heavily on Expert Parallelism, which generates intense inter-GPU communication. Consequently, state-of-the-art inference systems require high-bandwidth, P2P interconnects (e.g., NVLink) in datacenter GPUs to handle massive token routing, making deployment prohibitively expensive. Consumer GPUs offer comparable compute power at significantly lower cost, promising to democratize MoE inference for individuals and enable privacy-preserving local deployments. However, their bandwidth-limited (only weak PCIe bus bandwidth) and host-mediated interconnects (no P2P support) introduce severe communication bottlenecks.
We present CoMoE, a communication-efficient MoE inference system that resolves this mismatch through novel host-centric routing. Our key insight is that the unique communication topology provides the opportunity to elevate the host to an active routing hub, which can fundamentally reduce communication volume and eliminate global synchronization-induced stalls. Specifically, for token dispatch, we introduce host-backed token multicast to write shared tokens to the host exactly once, eliminating outbound transmission redundancy. For token combine, we propose a fine-grained, token-level aggregation mechanism using host staging buffers, which replaces rigid global synchronization and mitigates straggler effects. Evaluation on RTX 5090 GPUs shows that CoMoE improves inference throughput by up to 1.46x, approaching the performance of NVLink-capable A800 GPUs at only 23.4% of the hardware cost.