T-CCL: Resource Efficient and Performant Collective Communication using Tensor Memory Accelerator
This study addresses the excessive SM resource consumption of multi-GPU collective communication, which constrains computational concurrency. We propose the first resource-efficient communication scheme leveraging the Tensor Memory Accelerator (TMA). By offloading data movement and reduction operations to hardware and employing asynchronous pipelining for collective communication, our approach overcomes the thread-intensive bottlenecks of traditional libraries, significantly reducing SM overhead while maintaining high bandwidth and optimizing communication-computation overlap. Experimental results demonstrate that, compared to NCCL, the proposed method achieves up to 3.42× speedup with lower SM utilization, yields a 1.25× acceleration ratio when overlapped with GEMM operations, and improves end-to-end inference throughput by 1.31× when integrated into the vLLM backend.