GPU-Initiated Communication: Dissecting Down to the Bone

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the ambiguity in performance characteristics of GPU-initiated communication, where library overhead and hardware mechanisms are often conflated. By dissecting the GPU-NIC boundary communication path and constructing a minimal transport prototype, we systematically benchmark mainstream libraries—including RDMA, NVSHMEM, NCCL GIN, and DeepEP—on the H100 platform. Through decoupling software and hardware overhead, we quantify the latency cost of individual components, revealing that doorbell batching and queue parallelism are critical for achieving high throughput. Furthermore, our analysis demonstrates that shared queues induce significant latency spikes and that the submission path alone does not solely determine overall performance. This work provides fundamental mechanistic insights to guide communication optimization in AI clusters.
📝 Abstract
GPU-initiated communication lets GPU threads post RDMA operations directly to the NIC. It underpins NVSHMEM, NCCL GIN, and DeepEP, which serve the fine-grained, latency-critical communication of Mixture-of-Experts (MoE) models, yet its performance characteristics and optimizations remain scarcely documented beyond source code, and library comparisons fail to separate the costs of the hardware mechanism from those of the library around it. This paper dissects GPU-initiated communication at the GPU-NIC boundary. We first detail the GPU-side network path: queue placement, work-request construction, doorbell ordering, and completion semantics. We then introduce mini-gda and mini-proxy, minimal transports for the GPU and CPU-proxy submission paths, and measure them alongside NVSHMEM IBGDA, NCCL GIN, DeepEP, UCCL-EP, MSCCL++, and fabric-lib on NVIDIA H100, H200, B200, and GB200 platforms. A minimal GPU path issues an operation in 0.7 $μ$s and completes in 4.0 $μ$s; libraries add up to 4.6 $μ$s of issue time through queue management, memory ordering, and completion scope, and issue time scales with the SM clock. A tuned CPU proxy matches or beats the GPU path at idle, at the cost of a dedicated core whose operating state sets its latency and throughput. On either path, sharing a queue with bulk traffic raises latency by one to three orders of magnitude. Reaching the 260 M msg/s ceiling of our InfiniBand platform requires doorbell batching and queue parallelism, and both have resource costs: communication code can reduce GPU block residency even when unused, and all-to-all traffic loses 59% of its NIC message rate at about 3,000 active connections. The submission path alone therefore does not predict communication performance. Our experiment code and results are available at https://github.com/ParCoreLab/Dissecting-GPU-Communication-Experiments.
Problem

Research questions and friction points this paper is trying to address.

GPU-initiated communication
RDMA
Mixture-of-Experts
latency
NVSHMEM
Innovation

Methods, ideas, or system contributions that make the work stand out.

GPU-initiated communication
RDMA
dissection
minimal transport
Mixture-of-Experts
🔎 Similar Papers
2024-06-07International Symposium on High-Performance Computer ArchitectureCitations: 5