🤖 AI Summary
This study addresses the resource contention between collective communication and computation for GPU streaming multiprocessors (SMs) during distributed training of large models, which constrains overall training efficiency. To overcome this bottleneck, we propose HOCCL, a zero-SM collective communication framework that offloads communication entirely to DMA engines. Specifically, HOCCL incorporates three core components—a stream manager, a peer-to-peer executor, and a collective scheduler—to eliminate SM occupancy while preserving operator temporal ordering and maximizing bandwidth utilization. Experimental results demonstrate that HOCCL maintains near-peak communication performance while freeing approximately 10% of SM resources for computation. Consequently, end-to-end training throughput is improved by up to 5%, effectively resolving the communication-computation resource contention bottleneck in large-scale distributed training.
📝 Abstract
Large language model training involves massive computation on GPU streaming multiprocessors (SMs), the primary compute units of GPUs. Since SMs host specialized accelerators such as Tensor Cores, their efficient utilization is critical to training efficiency. Unfortunately, existing collective communication systems compete with computation for SMs, as they consume SMs for communication-related data movement and synchronization operations.
We observe that communication can, in principle, be driven by DMA engines, thereby eliminating SM involvement in communication. Based on this insight, we propose HOCCL, a zero-SM collective communication framework consisting of three components: a stream manager, a point-to-point (P2P) executor, and a collective scheduler. The stream manager preserves operator-level temporal ordering with other GPU kernels. The P2P executor enables zero-SM point-to-point communication, while the collective scheduler orchestrates P2P transfers to maximize bandwidth. Experiments show that HOCCL preserves near-peak communication performance, achieving within 3% of the state of the art on average, while eliminating communication occupancy on nearly 10% of total GPU SMs. By freeing SM resources for computation, HOCCL improves end-to-end training throughput by up to 5%.