optimize gpu performance

Design, build, or run profiling and benchmarking workflows and tools that measure GPU performance at kernel and system levels—breaking down kernel runtimes, instruction mix, memory behavior, warp-level stalls, and overall utilization. Analyze those measurements to tune kernels, memory layouts and transfers, occupancy/scheduler parameters (including CUDA settings), and other optimization techniques to improve throughput, latency, and resource utilization.

optimizegpuperformance

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.76
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$225K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the critical gap in existing performance analysis tools, which largely lack support for virtual GPUs (vGPUs) and thus struggle to diagnose performance bottlenecks in vGPU-accelerated applications. To tackle this limitation, the authors design and implement a novel vGPU performance profiling tool tailored for Intel GVT-g. The tool introduces, for the first time, a fine-grained software tracing mechanism coupled with runtime data collection to generate multidimensional performance metrics. These metrics are then presented through synchronized multi-view visualizations that capture comprehensive vGPU behavioral characteristics. By enabling detailed insight into vGPU execution dynamics, this approach significantly enhances observability and debugging efficiency in GPU-virtualized environments, effectively filling a longstanding void in vGPU performance analysis tooling.

debuggingGPU virtualizationperformance analysis

Privacy-Preserving Performance Profiling of In-The-Wild GPUs

Sep 25, 2025
IM
Ian McDougall
🏛️ University of Wisconsin-Madison | NVIDIA Research

This work addresses three key challenges in real-time GPU performance monitoring at planetary scale: user privacy leakage, runtime overhead, and difficulty in data attribution across massive deployments. We propose the first end-to-end privacy-preserving GPU performance profiling architecture that operates at planetary scale with zero performance overhead. Methodologically, we design a lightweight kernel-level monitoring agent—built upon NSYS/NCU—that integrates differential privacy, secure aggregation, and distributed sampling to enable application-agnostic, anonymized kernel-level behavior attribution and resource accounting. Evaluated on a simulated 100,000-GPU cluster, our system achieves complete telemetry collection while enabling precise application-level attribution for realistic deep learning workloads (e.g., TorchBench) with no third-party privacy leakage. The architecture delivers trustworthy, scalable hardware performance insights for chip design and systems optimization.

Collecting GPU performance data from real-world deployments at scaleEnabling profiling without performance slowdown for end usersPreserving user privacy while gathering hardware characteristics

A Performance Model for Warp Specialization Kernels

Jun 12, 2025
ZL
Zhengyang Liu
🏛️ University of Utah | NVIDIA

This work addresses the challenge of predicting execution performance for GPU warp-specialized kernels. We propose the first end-to-end performance model based on differential equations, jointly characterizing key factors including warp size, tiling dimensions, matrix dimensions, memory bandwidth, and thread divergence. The model is rigorously validated through both architectural analysis and empirical CUDA kernel measurements, augmented by a detailed bandwidth model. Its key innovation lies in the first application of differential equations to warp-level performance modeling, enabling quantitative characterization of the mapping between warp-level parallelism structures and performance bottlenecks. Experimental evaluation demonstrates a prediction error of less than 8.2%. The model supports compiler-driven auto-tuning and adaptive parameter configuration, achieving a 17% improvement in energy efficiency for sparse computation and GEMM workloads.

Develops a performance model for warp specialization kernelsOptimizes GPU applications via compiler and kernel tuningPredicts execution time using differential equations and validation

This work addresses the lack of transparency in NVIDIA’s closed-source user-space driver, which obscures the translation of CUDA API calls into hardware commands and impedes understanding of GPU behavior and performance attribution. The authors propose a novel approach that requires no modification to the proprietary driver, instead leveraging an open-source kernel driver, memory-mapped path instrumentation, and hardware watchpoints on the GPU’s doorbell registers to capture and reconstruct the complete low-level command stream with unprecedented accuracy. This methodology reveals the true DMA patterns and performance characteristics of CUDA data transfers and demonstrates that the low overhead of CUDA Graphs stems from their streamlined and efficient command submission mechanism. By significantly enhancing the interpretability of GPU runtime behavior, this approach establishes a new paradigm for middleware analysis and hardware-software co-design.

closed-source drivercommand streamCUDA

gpu_tracker: Python package for tracking and profiling GPU utilization in both desktop and high-performance computing environments

Apr 01, 2024
ED
Erik D. Huckvale
🏛️ University of Kentucky | Institute for Biomedical Informatics

Fine-grained, cross-platform real-time monitoring of GPU resources—particularly GPU memory peak usage and computational utilization—remains unsupported in Unix/Linux environments. Method: This paper introduces the first lightweight, dependency-free Python tool leveraging the NVIDIA Management Library (NVML) API. It employs multithreading and process-hooking techniques to enable low-overhead (average 0.3%) background sampling and precise peak capture of CPU/GPU utilization and system/GPU memory consumption. Contribution/Results: The tool unifies analysis across desktop and HPC environments with high accuracy (GPU memory peak error <2%). It enables job-level GPU resource profiling—the first such capability for fine-grained, runtime GPU characterization in HPC settings—thereby addressing a critical gap in production-grade GPU observability. The implementation is open-source and has been integrated into multiple scientific computing pipelines.

Monitors maximum RAM usage on motherboard and GPUProvides real-time resource profiling with minimal overheadTracks GPU and CPU utilization in HPC environments

Latest Papers

What's happening recently
View more

This work addresses the lack of automated, efficient methods for evaluating how closely existing benchmarks resemble real-world high-performance computing (HPC) applications in terms of hardware performance characteristics. The authors propose a novel performance similarity metric based on hardware usage patterns, introducing for the first time two distinct classes of computational kernels that exhibit similar performance behavior. They develop a scalable, automated evaluation framework that integrates performance feature analysis, kernel classification, and cross-platform (CPU/GPU) similarity assessment. The effectiveness and practicality of this approach are demonstrated by accurately matching computational kernels from the Kripke proxy application to those in the RAJA Performance Suite, thereby validating the method’s capability to identify functionally analogous kernels across diverse hardware architectures.

code representationcomputational kernelsHPC benchmarks

MT4G: A Tool for Reliable Auto-Discovery of NVIDIA and AMD GPU Compute and Memory Topologies

Nov 08, 2025
SV
Stepan Vanecek
🏛️ Technical University of Munich

GPU topology information—such as compute/memory hierarchy, cache sizes, interconnect bandwidths, and physical layout—is severely fragmented, incomplete, and vendor-locked, hindering performance modeling and optimization in HPC and AI systems. To address this, we propose the first cross-vendor (NVIDIA/AMD), open-source framework for automatic GPU hardware topology discovery. Our approach integrates CUDA/HIP APIs with over 50 customized microbenchmarks and employs statistical validation—including Kolmogorov–Smirnov tests—to infer non-programmable hardware attributes with high confidence. We validate the framework across ten mainstream GPUs, demonstrating broad compatibility and measurement accuracy. Furthermore, we integrate it into three critical workflows: analytical performance modeling, bottleneck analysis, and dynamic resource partitioning. This integration significantly enhances system-level hardware awareness and improves resource utilization efficiency.

Addressing vendor-specific limitations in GPU information accessAutomating GPU topology discovery for HPC and AI systemsProviding reliable identification of unavailable GPU topological attributes

This work addresses the limitation of existing GPU kernel generation benchmarks, which often diverge from real-world production workloads and thus fail to accurately evaluate the performance of large-model-generated kernels. To bridge this gap, we introduce Atrex-Bench—the first benchmark constructed from full-cluster, real inference traces with GPU time–weighted metrics—and present Atrex-Kernel-Agent, an optimization agent that integrates performance profiling with knowledge-driven strategies. Leveraging trace-driven sampling, iterative evaluation, an optimization discard mechanism, and a hierarchical knowledge base comprising 298 reference kernels and 244 documentation artifacts, our approach significantly enhances both correctness and efficiency of generated kernels. Experiments show that state-of-the-art coding agents achieve only ~10% of hardware roofline performance on Atrex-Bench, whereas our method eliminates reliance on PyTorch fallbacks and produces kernels that match or even surpass hand-tuned baselines.

kernel optimizationLLM-generated GPU kernelsperformance evaluation

This work proposes KernelPro, a closed-loop multi-agent system designed to automatically generate high-performance and energy-efficient GPU kernel code. By integrating large language models with hardware micro-benchmarking tools, KernelPro employs semantic feedback operators, a two-tier tool-calling architecture, a domain-adapted Monte Carlo Tree Search (MCTS) strategy, and direct CuTe source-code generation to jointly optimize for both performance and energy efficiency. Evaluated on KernelBench, KernelPro achieves up to a 5.30× speedup over baseline implementations. Furthermore, when applied to expert-optimized Mixture-of-Experts (MoE) kernels, it outperforms hand-tuned Triton kernels by 1.23× in performance while reducing measured energy consumption by 11.6%.

automatic code generationCUDAenergy efficiency

This study addresses the limitations of framework profilers lacking hardware counters, GPU profilers missing operator attribution, and the error-prone nature of manual correlation by proposing an automated hardware attribution pipeline. Methodologically, it integrates CUPTI, NVTX interval trees, and Inductor debug artifacts to design three complementary attribution paths alongside a call-order matching algorithm, incorporating deduplication and clock-locking mechanisms to ensure data consistency. Experimental results demonstrate that this approach achieves 95%–100% kernel runtime attribution for compilation workloads on Blackwell GPUs, enabling FX graph optimizations that yield up to a 2.24× speedup. This work provides a scalable solution for fine-grained performance bottleneck analysis.

GPU ProfilingHardware CountersOperator Profiling

Hot Scholars

IS

Ion Stoica

Professor of Computer Science, UC Berkeley
Cloud ComputingNetworkingDistributed SystemsBig Data
DZ

Danyang Zhuo

Duke University
Distributed SystemsNetworkingOperating Systems
EX

Enze Xie

NVIDIA Research, MMLab@HKU
computer visiongenerative AI
MG

Minyi Guo

IEEE Fellow, Chair Professor, Shanghai Jiao Tong University
Parallel ComputingCompiler OptimizationCloud ComputingNetworking
CK

Christos Kozyrakis

Stanford University
Computer ArchitectureComputer SystemsCloud Computing