🤖 AI Summary
This work addresses two critical challenges in ML training/inference: insufficient overlap between communication and computation, and severe interference within the GPU memory subsystem. We present the first systematic offloading of collective communication operations—including all-gather and all-to-all—to the DMA engine, implemented and evaluated on the AMD Instinct MI300X platform. To exploit DMA architectural characteristics, we propose a command scheduling mechanism and a lightweight synchronization protocol, effectively overcoming latency bottlenecks in small-data scenarios and mitigating performance degradation prevalent in prior offloading approaches. Experimental results show up to 16% higher throughput for large transfers and 32% lower power consumption; for small messages, performance gaps narrow significantly, with improvements of up to 7% in certain cases. This work is the first to comprehensively characterize the multi-dimensional trade-offs—across performance, energy efficiency, and synchronization overhead—of DMA-accelerated collective communication, establishing a foundational design paradigm and empirical basis for practical communication offloading in highly parallel AI systems.
📝 Abstract
Offloading machine learning (ML) communication collectives to direct memory access (DMA) engines has emerged as an interesting and low-cost solution to efficiently overlap computation and communication in inference and training. Doing so delivers superior concurrent performance by freeing up all GPU cores for computation and also lowers interference in the memory sub-system (caches). While DMA collectives show strong promise, prior works have only studied them in limited context (bandwidth-bound transfer sizes only, performance-only). To address this, we provide a comprehensive performance, power/energy and synchronization costs analysis of offloading ML communication collectives (all-gather, all-to-all) to DMA engines on state-of-the-art AMD Instinct MI300X GPUs. Our analysis reveals that, compared to the state-of-the-art RCCL communication collectives library, DMA collectives are at-par or better for large sizes (10s of MB to GB) in terms of both performance (16% better) and power (32% better). However, they significantly lag for latency-bound small sizes; 4.5X and 2.5X slower for all-gather and all-to-all, respectively. We provide a detailed latency breakdown of a DMA transfer and identify that DMA command scheduling and synchronization costs can limit DMA collective performance. To tackle this, we harness existing DMA architecture innovations, hitherto untapped, to build optimized DMA collectives and demonstrate their efficacy on real hardware. Our optimized implementations considerably close the performance gap for DMA collectives at smaller sizes (30% slower and 20% faster all-gather and all-to-all, respectively) and further improves performance (by 7%) and power savings at larger sizes (3-10%). Overall, this work represents a significant step toward making DMA collectives suitable for adoption in mainstream collective libraries.