🤖 AI Summary
This work addresses the lack of efficient and scalable concurrent queues on GPUs, which hinders task dispatching and resource utilization, thereby becoming a performance bottleneck in supercomputing systems. For the first time, it systematically explores the design space of GPU concurrent queues and proposes three novel structures—G-WFQ-YMC, G-LFQ, and G-WFQ—that collectively span correctness guarantees from lock-free to wait-free while preserving linearizability. Key optimizations include pre-allocated segments, wavefront-based batched fast paths, and 64-bit CAS packing of shared state to enhance concurrency control on GPUs. Experimental results demonstrate that the proposed queues substantially reduce core idle rates, improve parallel speedup and resource utilization, and achieve excellent scalability and throughput.
📝 Abstract
Concurrent queues can significantly impact supercomputing performance by being critical bottlenecks for task distribution, load balancing, and resource utilization. As HPC systems move beyond 10-million processor cores, the ability to rapidly move items between producer and consumer threads without excessive locking is essential for efficient queues, preventing idle cores, maximizing utilization, and achieving high parallel speedup. While concurrent queues are well studied on CPUs, they remain largely unexplored on modern GPUs, where SIMT execution, massive parallelism, and atomic contention reshape the design space. We present three linearizable GPU concurrent queues spanning from lock-free to wait-free guarantees: (1) G-WFQ-YMC, an adaptation of Yang and Mellor-Crummey's wait-free queue using preallocated segments; (2) G-LFQ, a bounded lock-free queue that uses wave-batched fast paths to maximize throughput; and (3) G-WFQ, a bounded wait-free queue that packs shared state into 64-bit compare-and-swap operations while preserving linearizability and bounded memory.