Score
Design, build, and analyze GPU compute kernels that keep threads active across multiple algorithmic steps (persistent threads), maintaining and updating mutable state in shared memory while minimizing global-memory traffic. Implement patterns such as iterative reductions, state-carrying clearing, and shared-memory atomics to cooperatively aggregate actions, perform per-block clearing, and reduce per-step critical-path depth so work and memory traffic scale independently of the step count.
This work addresses the performance degradation in large language model (LLM) inference on multi-chiplet NUMA GPUs, where non-uniform memory access and inter-chiplet communication induce kernel latency and poor memory locality. The study presents the first systematic classification of operand sharing patterns across workgroups in LLM kernels—categorized as global, partial, or private—and leverages memory trace analysis, workgroup-level access modeling, and cycle-accurate simulation to demonstrate that each pattern necessitates distinct data placement and scheduling strategies. Building on these insights, the authors propose a sub-group-aware co-scheduling mechanism coupled with an optimized data layout scheme, which significantly enhances execution efficiency and memory locality for LLM kernels on multi-chiplet GPU architectures.
SYCL programs on multi-GPU clusters suffer from high scheduling latency and substantial critical-path overhead due to implicit memory allocation, cache-coherence operations, and dependency analysis. Method: We propose the Instruction Graph—a novel intermediate representation that fully decouples scheduling from execution. Our approach integrates speculative scheduling, adaptive virtual-buffer memory allocation, and tight integration with the Celerity runtime, enabling fully concurrent scheduling of memory management, data transfers, MPI communication, and kernel launches while moving all scheduling analysis off the critical execution path. Contribution/Results: Evaluated on a production-scale 128-GPU cluster, our method achieves excellent strong scaling, drastically reduces multi-application scheduling latency, and drives critical-path overhead nearly to zero—thereby overcoming fundamental limitations of conventional static and blocking schedulers.
GPU programming under the weak memory model scoped-RC11 is prone to data races, barrier divergence, and assertion violations. Method: We propose GPUMC, a stateless model checker that formally encodes GPU-specific semantics—namely scoping, thread divergence, and hierarchical execution—into the scoped-RC11 model. GPUMC combines stateless depth-first search with constraint solving to enable full-path exploration and automatic error repair. Contributions/Results: (1) The first stateless model checking framework supporting native GPU concurrency semantics; (2) High-precision bug detection coupled with verifiable, automated fixes; (3) Comprehensive detection of known defects in real-world and benchmark GPU programs, with significantly lower time and memory overhead than state-of-the-art tools. GPUMC is the only tool to successfully verify several large-scale cases—others timed out.
In GPU programming, tension between fine-grained per-thread control and coarse-grained collective operations (e.g., Tensor Core instructions) undermines modularity and safety: collective primitives require coordinated execution across thread groups, yet encapsulated functions are invoked by individual threads. Method: We introduce Prism, a new language featuring *typed views*—a novel type-system mechanism that explicitly classifies thread behavior by control granularity (per-thread, group-wide, or global), enabling safe, modular abstraction of collective operations. Built upon the Bundl core calculus, Prism’s type-safe compiler supports precise, hardware-aware abstractions for accelerators like Tensor Cores. Contribution/Results: Evaluation shows Prism delivers strong type safety with zero runtime overhead, significantly improving GPU kernel correctness, maintainability, and developer productivity—without sacrificing performance.
To address the coexistence of performance instability and resource underutilization in multi-application GPU co-location, this paper introduces the first kernel-level, fine-grained resource interference quantification framework spanning multiple hardware layers—including compute units, L1/L2 caches, and memory bandwidth—overcoming the limitations of conventional coarse-grained utilization-based modeling. Leveraging micro-benchmarks, hardware performance counter sampling, kernel-level isolation experiments, and interference modeling, the framework enables reproducible characterization of interference behavior across critical subsystems. Based on this, we design a dynamic co-location scheduler with strict service-level objective (SLO) guarantees, achieving over 35% improvement in aggregate GPU utilization while maintaining quality-of-service requirements. This work establishes both theoretical foundations and empirical validation for predictable, high-performance GPU resource sharing.
This work addresses the inefficiency of existing fault-tolerance mechanisms for large language model (LLM) agents, which rely on full restarts or application-level checkpointing and struggle to recover critical GPU state—such as KV caches and scheduler metadata—efficiently. The authors propose a device-resident, persistent kernel implementation that enables fine-grained, low-overhead checkpointing and recovery directly within the GPU’s native execution context, thereby avoiding CPU bottlenecks. By injecting incremental checkpoint logic into the PTX/SASS layer via JIT compilation, and combining GPU module load interception, lock-free ring buffers, and CXL/host-memory logging, the approach achieves framework-agnostic, transparent fault tolerance. It supports efficient dirty-page detection and incremental snapshots of key structures like KV caches and adapter pages, substantially reducing work loss and recovery latency upon failure.
This work addresses the significant communication bottleneck in multi-GPU training caused by the serial execution of computation and communication. The authors propose a portable runtime mechanism that requires no modifications to vendor libraries or kernels. By dynamically controlling on-chip resource occupancy of compute kernels, elevating the scheduling priority of communication streams, and leveraging shared memory for compute footprint management and cross-GPU resource coordination, the approach effectively enables concurrent execution of computation and collective communication. Evaluated on NVIDIA A40, A100, H100, and AMD MI250X GPUs, the method reduces end-to-end training time by up to 25.5%.
Current GPU programming models lack expressiveness for chiplet-level locality and synchronization, leading to redundant memory accesses and poor cache utilization when executing memory-intensive workloads such as large language model (LLM) inference on multi-chiplet GPUs. This work proposes Fleet, the first multi-level task programming model that explicitly exposes the chiplet hierarchy. Fleet introduces a chiplet-task abstraction that binds computation and data to specific chiplets and integrates persistent kernels, cooperative weight tiling, and per-chiplet scheduling to enable L2 cache reuse and efficient coordinated execution. Evaluated on an AMD MI350 running Qwen3-8B, Fleet reduces decoding latency by 1.3–1.5× for small batches and cuts HBM traffic by up to 37% under large batches, significantly improving L2 hit rates and achieving overall speedups of 1.27–1.30×.
This study addresses the inefficiencies in GPU kernel benchmarking for LLM agents, where command-level exclusivity leads to poor resource utilization while shared execution compromises measurement precision. To resolve this trade-off, this work proposes a region-granularity exclusivity mechanism that enforces mutual exclusion solely within timing-critical regions, permitting concurrent execution elsewhere. By integrating runtime scheduling, process freezing, CPU core isolation, and persistent GPU context reuse, the proposed approach effectively balances high throughput with measurement fidelity. Experimental evaluations on both NVIDIA and AMD GPUs demonstrate up to a 3.4× improvement in throughput, with only 0.3% p95 time dilation observed for long-running kernels.
This work addresses the lack of efficient and scalable concurrent queues on GPUs, which hinders task dispatching and resource utilization, thereby becoming a performance bottleneck in supercomputing systems. For the first time, it systematically explores the design space of GPU concurrent queues and proposes three novel structures—G-WFQ-YMC, G-LFQ, and G-WFQ—that collectively span correctness guarantees from lock-free to wait-free while preserving linearizability. Key optimizations include pre-allocated segments, wavefront-based batched fast paths, and 64-bit CAS packing of shared state to enhance concurrency control on GPUs. Experimental results demonstrate that the proposed queues substantially reduce core idle rates, improve parallel speedup and resource utilization, and achieve excellent scalability and throughput.