Scalable Concurrent Queues for GPU

📅 2026-06-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the lack of efficient and scalable concurrent queues on GPUs, which hinders task dispatching and resource utilization, thereby becoming a performance bottleneck in supercomputing systems. For the first time, it systematically explores the design space of GPU concurrent queues and proposes three novel structures—G-WFQ-YMC, G-LFQ, and G-WFQ—that collectively span correctness guarantees from lock-free to wait-free while preserving linearizability. Key optimizations include pre-allocated segments, wavefront-based batched fast paths, and 64-bit CAS packing of shared state to enhance concurrency control on GPUs. Experimental results demonstrate that the proposed queues substantially reduce core idle rates, improve parallel speedup and resource utilization, and achieve excellent scalability and throughput.
📝 Abstract
Concurrent queues can significantly impact supercomputing performance by being critical bottlenecks for task distribution, load balancing, and resource utilization. As HPC systems move beyond 10-million processor cores, the ability to rapidly move items between producer and consumer threads without excessive locking is essential for efficient queues, preventing idle cores, maximizing utilization, and achieving high parallel speedup. While concurrent queues are well studied on CPUs, they remain largely unexplored on modern GPUs, where SIMT execution, massive parallelism, and atomic contention reshape the design space. We present three linearizable GPU concurrent queues spanning from lock-free to wait-free guarantees: (1) G-WFQ-YMC, an adaptation of Yang and Mellor-Crummey's wait-free queue using preallocated segments; (2) G-LFQ, a bounded lock-free queue that uses wave-batched fast paths to maximize throughput; and (3) G-WFQ, a bounded wait-free queue that packs shared state into 64-bit compare-and-swap operations while preserving linearizability and bounded memory.
Problem

Research questions and friction points this paper is trying to address.

concurrent queues
GPU
scalability
linearizability
high-performance computing
Innovation

Methods, ideas, or system contributions that make the work stand out.

GPU concurrent queues
wait-free
lock-free
linearizability
wave-batched
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
P
Pratheek Prakash Shetty
Department of ECE, Virginia Tech, Blacksburg, VA, USA
T
Thomas R. W. Scogland
Lawrence Livermore National Laboratory, Livermore, CA, USA
W
Wu-chun Feng
Department of CS and ECE, Virginia Tech, Blacksburg, VA, USA