Score
Designs, implements, and profiles CUDA/GPU kernels and their host orchestration to maximize throughput and efficiency on single- and multi‑GPU systems, including mapping computation to CUDA cores, configuring thread/block tiling, using shared memory and warp‑level primitives, and optimizing memory layout and transfers for coalescing and minimal contention. Applies hardware‑aware tuning — e.g., mixed‑precision/tensor‑core arithmetic, kernel fusion, batching and small‑batch strategies, segmented/warp reductions, multi‑stream and multi‑GPU scheduling, and platform‑specific parameterization — to accelerate training, inference, rendering and other GPU‑accelerated workloads while meeting deployment and performance objectives.
This work addresses the challenge of predicting execution performance for GPU warp-specialized kernels. We propose the first end-to-end performance model based on differential equations, jointly characterizing key factors including warp size, tiling dimensions, matrix dimensions, memory bandwidth, and thread divergence. The model is rigorously validated through both architectural analysis and empirical CUDA kernel measurements, augmented by a detailed bandwidth model. Its key innovation lies in the first application of differential equations to warp-level performance modeling, enabling quantitative characterization of the mapping between warp-level parallelism structures and performance bottlenecks. Experimental evaluation demonstrates a prediction error of less than 8.2%. The model supports compiler-driven auto-tuning and adaptive parameter configuration, achieving a 17% improvement in energy efficiency for sparse computation and GEMM workloads.
This work addresses the significant communication bottleneck in multi-GPU training caused by the serial execution of computation and communication. The authors propose a portable runtime mechanism that requires no modifications to vendor libraries or kernels. By dynamically controlling on-chip resource occupancy of compute kernels, elevating the scheduling priority of communication streams, and leveraging shared memory for compute footprint management and cross-GPU resource coordination, the approach effectively enables concurrent execution of computation and collective communication. Evaluated on NVIDIA A40, A100, H100, and AMD MI250X GPUs, the method reduces end-to-end training time by up to 25.5%.
Frequent fine-grained kernel launches on GPUs incur substantial launch overhead, severely limiting performance in scientific computing. To address this, we propose a synergistic optimization combining iterative batching and CUDA Graph unrolling: multiple iterations are grouped into batches and statically unrolled into a single CUDA Graph, thereby eliminating redundant kernel launch overhead. We further introduce the first platform-agnostic criterion for selecting the optimal batch size and develop a generalizable analytical performance model. Evaluated on skeleton applications, our approach achieves over 1.4× speedup. It demonstrates significant and robust performance improvements across real-world iterative GPU applications—including Hotspot, Hotspot3D, and an FDTD-based Maxwell solver—without requiring application-specific tuning. This work establishes a general, analytically tractable, low-overhead execution paradigm for iterative GPU computations.
This work addresses the throughput limitations of GPU systems caused by host-device synchronization latency and kernel scheduling overhead, which hinder efficient utilization of compute cores and copy engines. The authors propose a CUDA runtime framework tailored for task-parallel pipelining, which innovatively integrates multi-stream scheduling, event-chain triggering, work stealing, and stream-level buffer management. This design ensures memory safety across concurrent tasks while substantially reducing kernel launch intervals and synchronization overhead. Implemented using CUDA Graphs, the framework achieves 1.15–1.44× speedup over state-of-the-art baselines on real-world workloads and reduces scheduling overhead by 18%–54%.
Existing LLM-driven approaches for GPU kernel optimization are largely confined to machine learning scenarios, lacking cross-domain generality and a systematic evaluation benchmark. To address this gap, this work proposes CUDAMaster—a hardware-aware, multi-agent automated optimization framework that integrates performance profiling with an automated compilation toolchain to generate CUDA kernels across diverse workloads under FP32 and BF16 precision. Furthermore, we introduce MSKernelBench, the first comprehensive benchmark encompassing algebraic operations, LLM operators, sparse matrix computations, and scientific computing. Experimental results demonstrate that CUDAMaster consistently outperforms Astra by an average of approximately 35% across most operators, with certain kernels matching or even surpassing the performance of the proprietary cuBLAS library.
This work addresses the challenge of optimizing and analyzing GPU kernels for deep convolution in cloud environments lacking hardware performance counters. The authors propose a counter-free performance analysis framework that integrates execution path decomposition with memory traffic modeling to systematically optimize the forward, input gradient, and weight gradient computation paths. Leveraging CUDA event timing, effective bandwidth estimation, and the roofline model, the approach employs optimization strategies such as shared memory tiling and warp-level partitioning. The resulting warp-tiled kernels achieve a 3.26× speedup over naive implementations, yielding a 1.29× end-to-end training acceleration. This significantly enhances the efficiency of critical computational paths and provides reproducible, architecture-level performance insights for resource-constrained cloud settings.
This work systematically investigates the trade-offs between performance and programmability in CPU-GPU cooperative scheduling across discrete and unified memory architectures, with a focus on sparse conjugate gradient computations. Evaluations are conducted on both the NVIDIA GH200 Superchip—a platform featuring a unified memory architecture—and the discrete H100 PCIe system, comparing three memory management paradigms: explicit data copies, managed memory, and mapped memory. The study reveals that the GH200’s fused architecture substantially enhances the practicality of managed memory, enabling diverse hybrid task-partitioning strategies to achieve both high performance and programming simplicity. These findings underscore the significant impact of underlying memory architecture on the efficacy of cooperative scheduling approaches.
This work proposes KernelPro, a closed-loop multi-agent system designed to automatically generate high-performance and energy-efficient GPU kernel code. By integrating large language models with hardware micro-benchmarking tools, KernelPro employs semantic feedback operators, a two-tier tool-calling architecture, a domain-adapted Monte Carlo Tree Search (MCTS) strategy, and direct CuTe source-code generation to jointly optimize for both performance and energy efficiency. Evaluated on KernelBench, KernelPro achieves up to a 5.30× speedup over baseline implementations. Furthermore, when applied to expert-optimized Mixture-of-Experts (MoE) kernels, it outperforms hand-tuned Triton kernels by 1.23× in performance while reducing measured energy consumption by 11.6%.
This work addresses the performance degradation in large language model (LLM) inference on multi-chiplet NUMA GPUs, where non-uniform memory access and inter-chiplet communication induce kernel latency and poor memory locality. The study presents the first systematic classification of operand sharing patterns across workgroups in LLM kernels—categorized as global, partial, or private—and leverages memory trace analysis, workgroup-level access modeling, and cycle-accurate simulation to demonstrate that each pattern necessitates distinct data placement and scheduling strategies. Building on these insights, the authors propose a sub-group-aware co-scheduling mechanism coupled with an optimized data layout scheme, which significantly enhances execution efficiency and memory locality for LLM kernels on multi-chiplet GPU architectures.