gpus

Designs, implements, and analyzes GPU hardware and software systems—including accelerator microarchitecture, drivers, runtimes, and compute kernels—to optimize parallel throughput, memory hierarchy behavior, scheduling, and energy efficiency. Work includes developing and tuning GPU-accelerated algorithms, profiling and optimizing kernels, integrating GPUs into larger systems, and evaluating performance, reliability, and power tradeoffs.

gpus

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.95
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$215K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

FPGA or GPU? Analyzing Comparative Research for Application-Specific Guidance

Mar 22, 2025
AA
Arnab A Purkayastha
🏛️ Western New England University | The Citadel

Existing FPGA–GPU comparative studies predominantly focus on raw performance metrics and lack domain-specific guidance for accelerator selection. Method: This paper proposes an application-oriented, fine-grained comparative framework that systematically synthesizes over 100 studies, conducting cross-domain (e.g., AI, HPC, network processing, scientific computing) classification and cross-evaluation along three dimensions: performance, energy efficiency, and programmability. Contribution/Results: The study innovatively establishes the first empirically grounded applicability boundaries for FPGAs and GPUs: FPGAs excel in low-latency, high-throughput customized pipelines and energy-constrained scenarios; GPUs are superior for massively parallel, computation-intensive workloads with stable algorithms. The resulting actionable decision-making guide enables researchers and engineers to select hardware accelerators based on domain-specific requirements, thereby bridging the gap between architectural characteristics and real-world application needs.

Analyzing FPGA vs GPU performance for application-specific guidanceBridging research gap on ideal accelerators for domain-specific applicationsProviding actionable recommendations for hardware accelerator selection

Addressing the challenge of kernel optimization on next-generation GPUs (e.g., AMD MI300), where scarce documentation and limited expert knowledge hinder manual tuning, this paper proposes an LLM-driven, multi-stage evolutionary agent framework. The framework operates without human priors, autonomously generating optimization hypotheses, evolving CUDA/HIP code variants, scheduling experiments, and iterating closed-loop based solely on runtime performance feedback. It integrates general GPU optimization knowledge from literature, architecture-specific adaptation mechanisms, and external evaluation systems to enable end-to-end automated tuning. Experiments demonstrate its feasibility in complex heterogeneous hardware environments, substantially reducing reliance on domain experts. To our knowledge, this is the first work to deeply apply LLMs to low-level, system-level performance optimization—establishing a novel paradigm for efficient programming of emerging accelerator architectures.

Addressing challenges in newer or less-documented GPU architecturesAutomating GPU kernel optimization for high performanceLeveraging LLMs to reduce need for domain expertise

Data-Driven Analysis to Understand GPU Hardware Resource Usage of Optimizations

Aug 19, 2024
TZ
Tanzima Z. Islam
🏛️ Texas State University

To address low per-GPU utilization and suboptimal hardware return-on-investment in heterogeneous multi-GPU systems, this paper proposes a data-driven analytical framework that establishes, for the first time, interpretable correlations between optimization strategies and GPU resource usage patterns. Our method integrates hardware performance counter profiling, multi-objective correlation modeling, and scientific proxy application benchmarks to construct a multidimensional metric suite characterizing application-device interaction behaviors. Unlike prior work—which focuses primarily on performance gains—our approach systematically uncovers the underlying mechanisms by which optimizations affect resource occupancy and utilization. Experimental evaluation on proxy applications demonstrates a 29.6% reduction in execution time, a 5.3% increase in average GPU utilization, and a 26.5% decrease in power consumption. These results establish a novel paradigm for resource-efficient, co-optimized heterogeneous accelerator systems.

Analyzes GPU hardware resource usage impact on utilization and performance.Develops multi-objective metric to optimize application-device interactions.Identifies optimization opportunities for scientific applications based on resource usage.

This work proposes KernelPro, a closed-loop multi-agent system designed to automatically generate high-performance and energy-efficient GPU kernel code. By integrating large language models with hardware micro-benchmarking tools, KernelPro employs semantic feedback operators, a two-tier tool-calling architecture, a domain-adapted Monte Carlo Tree Search (MCTS) strategy, and direct CuTe source-code generation to jointly optimize for both performance and energy efficiency. Evaluated on KernelBench, KernelPro achieves up to a 5.30× speedup over baseline implementations. Furthermore, when applied to expert-optimized Mixture-of-Experts (MoE) kernels, it outperforms hand-tuned Triton kernels by 1.23× in performance while reducing measured energy consumption by 11.6%.

automatic code generationCUDAenergy efficiency

SwizzlePerf: Hardware-Aware LLMs for GPU Kernel Performance Optimization

Aug 27, 2025
AT
Arya Tschand
🏛️ Harvard University | AMD | Stanford University

Current search-based LLM approaches lack hardware awareness, hindering near-optimal GPU kernel performance optimization. This paper introduces the first hardware-aware LLM framework for automated spatial optimization targeting disaggregated architectures. Our method integrates memory access pattern analysis, fine-grained architectural modeling, and curated historical performance logs with real-time feedback to guide LLMs in collaboratively generating optimal swizzling strategies. Evaluated on ten representative ML and scientific computing kernels, our approach achieves up to 2.06× speedup on nine kernels—accompanied by a 70% improvement in L2 cache hit rate—and reduces GEMM optimization time from two weeks to under five minutes. The framework significantly enhances both hardware-software co-optimization efficiency and generalizability across diverse workloads and architectures.

Addressing inefficiency in search-based runtime optimization methodsAutomating spatial optimizations for disaggregated GPU architecturesOptimizing GPU kernel performance using hardware-aware LLMs

Latest Papers

What's happening recently
View more

This work addresses the unpredictability of GPU kernel execution times caused by inter-kernel data dependencies and resource contention, which undermines real-time guarantees. Focusing on DAG-structured GPU tasks, the paper proposes a scheduling approach that does not rely on kernel priority assumptions. By explicitly modeling kernel dependencies and co-scheduling kernel-level parallelism, the method derives tight and safe worst-case execution time (WCET) bounds. Built upon standard CUDA APIs, the approach requires only DAG task modeling, parallel kernel scheduling, and response time analysis, without additional hardware or software support. Experimental results on both synthetic and real-world benchmarks demonstrate that the proposed method reduces WCET and measured execution time by up to 32.8% and 21.3%, respectively, compared to existing techniques.

DAGdata dependenciesGPU tasks

This work addresses the critical gap in existing performance analysis tools, which largely lack support for virtual GPUs (vGPUs) and thus struggle to diagnose performance bottlenecks in vGPU-accelerated applications. To tackle this limitation, the authors design and implement a novel vGPU performance profiling tool tailored for Intel GVT-g. The tool introduces, for the first time, a fine-grained software tracing mechanism coupled with runtime data collection to generate multidimensional performance metrics. These metrics are then presented through synchronized multi-view visualizations that capture comprehensive vGPU behavioral characteristics. By enabling detailed insight into vGPU execution dynamics, this approach significantly enhances observability and debugging efficiency in GPU-virtualized environments, effectively filling a longstanding void in vGPU performance analysis tooling.

debuggingGPU virtualizationperformance analysis

Existing LLM-driven approaches for GPU kernel optimization are largely confined to machine learning scenarios, lacking cross-domain generality and a systematic evaluation benchmark. To address this gap, this work proposes CUDAMaster—a hardware-aware, multi-agent automated optimization framework that integrates performance profiling with an automated compilation toolchain to generate CUDA kernels across diverse workloads under FP32 and BF16 precision. Furthermore, we introduce MSKernelBench, the first comprehensive benchmark encompassing algebraic operations, LLM operators, sparse matrix computations, and scientific computing. Experimental results demonstrate that CUDAMaster consistently outperforms Astra by an average of approximately 35% across most operators, with certain kernels matching or even surpassing the performance of the proprietary cuBLAS library.

benchmarkCUDA kernel optimizationLLM