Optimizing ML Concurrent Computation and Communication with GPU DMA Engines

📅 2024-12-18
🏛️ arXiv.org
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Computation-communication concurrency (C3) for ML training/inference on GPUs suffers severe performance degradation—achieving only 21% of the ideal speedup—primarily due to strong interference between compute and memory bandwidth resources. This work introduces the first GPU-specific DMA-engine-based communication offloading mechanism and proposes ConCCL, a lightweight concurrent communication primitive. We jointly mitigate resource contention via kernel scheduling priority control and dynamic Streaming Multiprocessor (SM) partitioning. Leveraging DMA programming, fine-grained resource isolation, customized collective implementations, and heuristic runtime optimizations, our approach improves the effective C3 speedup to 72% on average, with end-to-end acceleration reaching up to 1.67×. This work establishes a novel architectural paradigm for efficient C3 on GPUs and significantly extends the frontier of heterogeneous parallel optimization.

Technology Category

Machine Learning: Hardware-aware MLConstraint Satisfaction and Optimization: Distributed CSP/OptimizationData Mining & Knowledge Management: Scalability, Parallel & Distributed Systems

Application Category

Economics, Online Markets and Human Computation: Architectures and workflows that use LLMs for crowd workSearch and Retrieval-Augmented AI: Efficiency and scalability of Web search enginesGraph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphs
📝 Abstract
Concurrent computation and communication (C3) is a pervasive paradigm in ML and other domains, making its performance optimization crucial. In this paper, we carefully characterize C3 in ML on GPUs, which are most widely deployed for ML training and inference. We observe that while C3 leads to performance uplifts, the uplifts are far lower than ideal speedups (serial computation and communication versus maximum of computation or communication; all times from isolated executions). That is, C3 on average achieves only 21% of ideal speedup. This is so, due to known challenges of compute and memory interference between concurrent GPU kernels (that is, sharing of GPU's compute units, caches and HBM). To attain better performance for C3, first, we evaluate dual strategies of schedule prioritization and careful resource partitioning of compute units on GPUs to push performance attained with C3 (on average 42% of ideal speedup). We also provide heuristics that can guide a runtime while employing these strategies. To further enhance C3 performance, we propose to mitigate C3 interference by offloading communication tasks to the GPU's DMA engines. To this end, we build concurrent communication collectives (ConCCL) proof-of-concepts that harness DMA engines for communication. We show how ConCCL considerably closes the gap between realized and ideal speedup for C3 (on average 72% of ideal speedup is realized, up to 1.67x speedup). Overall, our work makes a strong case for GPU DMA engine advancements to better support C3 on GPUs.
Problem

Research questions and friction points this paper is trying to address.

Optimizing concurrent computation and communication in ML on GPUs
Addressing performance gaps caused by GPU kernel interference
Enhancing C3 performance using GPU DMA engines for communication
Innovation

Methods, ideas, or system contributions that make the work stand out.

Optimize C3 with GPU DMA engines
Prioritize schedule and partition resources
Build ConCCL for concurrent communication
🔎 Similar Papers
2024-06-07International Symposium on High-Performance Computer ArchitectureCitations: 5
Advanced Micro Devices | Massey University
A
Anirudha Agrawal
Advanced Micro Devices, Inc.
Shaizeen Aga
Shaizeen Aga
AMD Research
Near-data processingSecure hardwareParallel Computer Architecture
S
Suchita Pati
Advanced Micro Devices, Inc.
M
Mahzabeen Islam
Advanced Micro Devices, Inc.