gpu computing

Designs, implements, and optimizes compute kernels, data pipelines, and algorithms to run efficiently on graphics processing units (GPUs), including writing GPU code (e.g., CUDA, OpenCL, or similar), managing memory hierarchy and host-device transfers, parallelization, synchronization, and precision choices. Builds and evaluates GPU-accelerated components or systems and analyzes their performance, resource utilization, scalability, and numerical behavior using profiling and benchmarking tools.

gpucomputing

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.22
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$203K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the critical gap in existing performance analysis tools, which largely lack support for virtual GPUs (vGPUs) and thus struggle to diagnose performance bottlenecks in vGPU-accelerated applications. To tackle this limitation, the authors design and implement a novel vGPU performance profiling tool tailored for Intel GVT-g. The tool introduces, for the first time, a fine-grained software tracing mechanism coupled with runtime data collection to generate multidimensional performance metrics. These metrics are then presented through synchronized multi-view visualizations that capture comprehensive vGPU behavioral characteristics. By enabling detailed insight into vGPU execution dynamics, this approach significantly enhances observability and debugging efficiency in GPU-virtualized environments, effectively filling a longstanding void in vGPU performance analysis tooling.

debuggingGPU virtualizationperformance analysis

Data-Driven Analysis to Understand GPU Hardware Resource Usage of Optimizations

Aug 19, 2024
TZ
Tanzima Z. Islam
🏛️ Texas State University

To address low per-GPU utilization and suboptimal hardware return-on-investment in heterogeneous multi-GPU systems, this paper proposes a data-driven analytical framework that establishes, for the first time, interpretable correlations between optimization strategies and GPU resource usage patterns. Our method integrates hardware performance counter profiling, multi-objective correlation modeling, and scientific proxy application benchmarks to construct a multidimensional metric suite characterizing application-device interaction behaviors. Unlike prior work—which focuses primarily on performance gains—our approach systematically uncovers the underlying mechanisms by which optimizations affect resource occupancy and utilization. Experimental evaluation on proxy applications demonstrates a 29.6% reduction in execution time, a 5.3% increase in average GPU utilization, and a 26.5% decrease in power consumption. These results establish a novel paradigm for resource-efficient, co-optimized heterogeneous accelerator systems.

Analyzes GPU hardware resource usage impact on utilization and performance.Develops multi-objective metric to optimize application-device interactions.Identifies optimization opportunities for scientific applications based on resource usage.

Addressing the challenge of kernel optimization on next-generation GPUs (e.g., AMD MI300), where scarce documentation and limited expert knowledge hinder manual tuning, this paper proposes an LLM-driven, multi-stage evolutionary agent framework. The framework operates without human priors, autonomously generating optimization hypotheses, evolving CUDA/HIP code variants, scheduling experiments, and iterating closed-loop based solely on runtime performance feedback. It integrates general GPU optimization knowledge from literature, architecture-specific adaptation mechanisms, and external evaluation systems to enable end-to-end automated tuning. Experiments demonstrate its feasibility in complex heterogeneous hardware environments, substantially reducing reliance on domain experts. To our knowledge, this is the first work to deeply apply LLMs to low-level, system-level performance optimization—establishing a novel paradigm for efficient programming of emerging accelerator architectures.

Addressing challenges in newer or less-documented GPU architecturesAutomating GPU kernel optimization for high performanceLeveraging LLMs to reduce need for domain expertise

Taking GPU Programming Models to Task for Performance Portability

Feb 14, 2024
JH
Joshua H. Davis
🏛️ University of Maryland | Lawrence Livermore National Laboratory

Assessing performance portability across GPU programming models on heterogeneous NVIDIA and AMD hardware remains challenging due to fragmented benchmarks and irreproducible evaluation methodologies. Method: We conduct a systematic, cross-platform evaluation of seven programming models—CUDA, HIP, Kokkos, RAJA, OpenMP, OpenACC, and SYCL—using five interdisciplinary proxy applications. Leveraging a Spack-based automation framework, we ensure fully reproducible build, deployment, and benchmarking across real multi-vendor GPU systems. Contribution/Results: Our empirical study quantifies both performance consistency and migration overhead for each model. HIP and SYCL achieve superior performance on AMD GPUs; CUDA remains dominant on NVIDIA hardware; Kokkos and RAJA deliver balanced portability with moderate performance; OpenMP and OpenACC exhibit significant cross-platform performance degradation. This work provides the first unified, vendor-agnostic assessment of GPU programming models’ performance portability and establishes a rigorous, reproducible methodology to guide architecture selection for high-performance scientific software.

Analyzing underperformance causes and providing optimizationsAssessing consistency on NVIDIA and AMD GPU architecturesEvaluating performance portability across GPU programming models

Latest Papers

What's happening recently
View more

This work proposes KernelPro, a closed-loop multi-agent system designed to automatically generate high-performance and energy-efficient GPU kernel code. By integrating large language models with hardware micro-benchmarking tools, KernelPro employs semantic feedback operators, a two-tier tool-calling architecture, a domain-adapted Monte Carlo Tree Search (MCTS) strategy, and direct CuTe source-code generation to jointly optimize for both performance and energy efficiency. Evaluated on KernelBench, KernelPro achieves up to a 5.30× speedup over baseline implementations. Furthermore, when applied to expert-optimized Mixture-of-Experts (MoE) kernels, it outperforms hand-tuned Triton kernels by 1.23× in performance while reducing measured energy consumption by 11.6%.

automatic code generationCUDAenergy efficiency

Existing LLM-driven approaches for GPU kernel optimization are largely confined to machine learning scenarios, lacking cross-domain generality and a systematic evaluation benchmark. To address this gap, this work proposes CUDAMaster—a hardware-aware, multi-agent automated optimization framework that integrates performance profiling with an automated compilation toolchain to generate CUDA kernels across diverse workloads under FP32 and BF16 precision. Furthermore, we introduce MSKernelBench, the first comprehensive benchmark encompassing algebraic operations, LLM operators, sparse matrix computations, and scientific computing. Experimental results demonstrate that CUDAMaster consistently outperforms Astra by an average of approximately 35% across most operators, with certain kernels matching or even surpassing the performance of the proprietary cuBLAS library.

benchmarkCUDA kernel optimizationLLM

This work addresses the lack of transparency in NVIDIA’s closed-source user-space driver, which obscures the translation of CUDA API calls into hardware commands and impedes understanding of GPU behavior and performance attribution. The authors propose a novel approach that requires no modification to the proprietary driver, instead leveraging an open-source kernel driver, memory-mapped path instrumentation, and hardware watchpoints on the GPU’s doorbell registers to capture and reconstruct the complete low-level command stream with unprecedented accuracy. This methodology reveals the true DMA patterns and performance characteristics of CUDA data transfers and demonstrates that the low overhead of CUDA Graphs stems from their streamlined and efficient command submission mechanism. By significantly enhancing the interpretability of GPU runtime behavior, this approach establishes a new paradigm for middleware analysis and hardware-software co-design.

closed-source drivercommand streamCUDA

Hot Scholars

KD

Karthik Dantu

Associate Professor, University at Buffalo
RoboticsMobile SystemsSensing SystemsEdge Computing
WZ

Weiguang Zhao

Univeristy of Liverpool, PhD Candidate
3D VisionEmbodied AIOpen World
KH

Kaizhu Huang

Professor, Duke Kunshan University
Generalization & RobustnessStatistical Learning ThoeryTrustworthy AI
NV

Nicholas Vining

Sr. Research Scientist, NVIDIA
Geometric Mesh ProcessingReal-Time RenderingGame DevelopmentComputer Graphics
AS

Alla Sheffer

Professor, Computer Science, University of British Columbia, Canada
Computer Graphicsgraphicsgeometry processinggeometric modeling