nvidia gpus

Designs, builds, and evaluates systems and software that use NVIDIA GPUs, including writing and optimizing CUDA/OpenCL kernels, managing drivers and libraries, and integrating multi‑GPU configurations. Involves profiling and benchmarking, tuning memory and compute concurrency, and diagnosing hardware and software performance, scalability, and power/thermal behavior for GPU‑accelerated applications.

nvidiagpus

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.52
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$215K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the lack of transparency in NVIDIA’s closed-source user-space driver, which obscures the translation of CUDA API calls into hardware commands and impedes understanding of GPU behavior and performance attribution. The authors propose a novel approach that requires no modification to the proprietary driver, instead leveraging an open-source kernel driver, memory-mapped path instrumentation, and hardware watchpoints on the GPU’s doorbell registers to capture and reconstruct the complete low-level command stream with unprecedented accuracy. This methodology reveals the true DMA patterns and performance characteristics of CUDA data transfers and demonstrates that the low overhead of CUDA Graphs stems from their streamlined and efficient command submission mechanism. By significantly enhancing the interpretability of GPU runtime behavior, this approach establishes a new paradigm for middleware analysis and hardware-software co-design.

closed-source drivercommand streamCUDA

This work addresses the critical gap in existing performance analysis tools, which largely lack support for virtual GPUs (vGPUs) and thus struggle to diagnose performance bottlenecks in vGPU-accelerated applications. To tackle this limitation, the authors design and implement a novel vGPU performance profiling tool tailored for Intel GVT-g. The tool introduces, for the first time, a fine-grained software tracing mechanism coupled with runtime data collection to generate multidimensional performance metrics. These metrics are then presented through synchronized multi-view visualizations that capture comprehensive vGPU behavioral characteristics. By enabling detailed insight into vGPU execution dynamics, this approach significantly enhances observability and debugging efficiency in GPU-virtualized environments, effectively filling a longstanding void in vGPU performance analysis tooling.

debuggingGPU virtualizationperformance analysis

MT4G: A Tool for Reliable Auto-Discovery of NVIDIA and AMD GPU Compute and Memory Topologies

Nov 08, 2025
SV
Stepan Vanecek
🏛️ Technical University of Munich

GPU topology information—such as compute/memory hierarchy, cache sizes, interconnect bandwidths, and physical layout—is severely fragmented, incomplete, and vendor-locked, hindering performance modeling and optimization in HPC and AI systems. To address this, we propose the first cross-vendor (NVIDIA/AMD), open-source framework for automatic GPU hardware topology discovery. Our approach integrates CUDA/HIP APIs with over 50 customized microbenchmarks and employs statistical validation—including Kolmogorov–Smirnov tests—to infer non-programmable hardware attributes with high confidence. We validate the framework across ten mainstream GPUs, demonstrating broad compatibility and measurement accuracy. Furthermore, we integrate it into three critical workflows: analytical performance modeling, bottleneck analysis, and dynamic resource partitioning. This integration significantly enhances system-level hardware awareness and improves resource utilization efficiency.

Addressing vendor-specific limitations in GPU information accessAutomating GPU topology discovery for HPC and AI systemsProviding reliable identification of unavailable GPU topological attributes

Deep Learning and Machine Learning with GPGPU and CUDA: Unlocking the Power of Parallel Computing

Oct 08, 2024
ML
Ming Li
🏛️ Georgia Institute of Technology | Indiana University | Xi'an Jiaotong-Liverpool University | University of Hawaii | Kyoto University | AppCubic | Rutgers University | University of Wisconsin-Madison | Purdue University | National Taiwan Normal University

This study addresses the GPU computational efficiency bottleneck in deep and machine learning. Methodologically, it proposes a task-aware GPU parallel architecture adaptation framework that systematically integrates CUDA stream-based concurrency, dynamic parallelism, and heterogeneous hardware (FPGA/TPU/ASIC) co-selection—implemented via deep integration into PyTorch, TensorFlow, and XGBoost. Its core contribution lies in establishing a transferable GPU optimization methodology, transcending model- or library-specific tuning. Experimental evaluation demonstrates 3–8× speedup across representative training and inference workloads. Furthermore, the authors open-source a modular, well-documented GPU optimization practice guide, substantially lowering the barrier to entry for AI practitioners seeking parallelization optimizations.

Exploring GPU architectures for efficient parallel computing in machine learningOptimizing deep learning algorithms using CUDA and GPGPU acceleration techniquesSelecting appropriate parallel hardware for specific computational tasks and frameworks

Taking GPU Programming Models to Task for Performance Portability

Feb 14, 2024
JH
Joshua H. Davis
🏛️ University of Maryland | Lawrence Livermore National Laboratory

Assessing performance portability across GPU programming models on heterogeneous NVIDIA and AMD hardware remains challenging due to fragmented benchmarks and irreproducible evaluation methodologies. Method: We conduct a systematic, cross-platform evaluation of seven programming models—CUDA, HIP, Kokkos, RAJA, OpenMP, OpenACC, and SYCL—using five interdisciplinary proxy applications. Leveraging a Spack-based automation framework, we ensure fully reproducible build, deployment, and benchmarking across real multi-vendor GPU systems. Contribution/Results: Our empirical study quantifies both performance consistency and migration overhead for each model. HIP and SYCL achieve superior performance on AMD GPUs; CUDA remains dominant on NVIDIA hardware; Kokkos and RAJA deliver balanced portability with moderate performance; OpenMP and OpenACC exhibit significant cross-platform performance degradation. This work provides the first unified, vendor-agnostic assessment of GPU programming models’ performance portability and establishes a rigorous, reproducible methodology to guide architecture selection for high-performance scientific software.

Analyzing underperformance causes and providing optimizationsAssessing consistency on NVIDIA and AMD GPU architecturesEvaluating performance portability across GPU programming models

Latest Papers

What's happening recently
View more

Privacy-Preserving Performance Profiling of In-The-Wild GPUs

Sep 25, 2025
IM
Ian McDougall
🏛️ University of Wisconsin-Madison | NVIDIA Research

This work addresses three key challenges in real-time GPU performance monitoring at planetary scale: user privacy leakage, runtime overhead, and difficulty in data attribution across massive deployments. We propose the first end-to-end privacy-preserving GPU performance profiling architecture that operates at planetary scale with zero performance overhead. Methodologically, we design a lightweight kernel-level monitoring agent—built upon NSYS/NCU—that integrates differential privacy, secure aggregation, and distributed sampling to enable application-agnostic, anonymized kernel-level behavior attribution and resource accounting. Evaluated on a simulated 100,000-GPU cluster, our system achieves complete telemetry collection while enabling precise application-level attribution for realistic deep learning workloads (e.g., TorchBench) with no third-party privacy leakage. The architecture delivers trustworthy, scalable hardware performance insights for chip design and systems optimization.

Collecting GPU performance data from real-world deployments at scaleEnabling profiling without performance slowdown for end usersPreserving user privacy while gathering hardware characteristics

Existing LLM-driven approaches for GPU kernel optimization are largely confined to machine learning scenarios, lacking cross-domain generality and a systematic evaluation benchmark. To address this gap, this work proposes CUDAMaster—a hardware-aware, multi-agent automated optimization framework that integrates performance profiling with an automated compilation toolchain to generate CUDA kernels across diverse workloads under FP32 and BF16 precision. Furthermore, we introduce MSKernelBench, the first comprehensive benchmark encompassing algebraic operations, LLM operators, sparse matrix computations, and scientific computing. Experimental results demonstrate that CUDAMaster consistently outperforms Astra by an average of approximately 35% across most operators, with certain kernels matching or even surpassing the performance of the proprietary cuBLAS library.

benchmarkCUDA kernel optimizationLLM

This work proposes KernelPro, a closed-loop multi-agent system designed to automatically generate high-performance and energy-efficient GPU kernel code. By integrating large language models with hardware micro-benchmarking tools, KernelPro employs semantic feedback operators, a two-tier tool-calling architecture, a domain-adapted Monte Carlo Tree Search (MCTS) strategy, and direct CuTe source-code generation to jointly optimize for both performance and energy efficiency. Evaluated on KernelBench, KernelPro achieves up to a 5.30× speedup over baseline implementations. Furthermore, when applied to expert-optimized Mixture-of-Experts (MoE) kernels, it outperforms hand-tuned Triton kernels by 1.23× in performance while reducing measured energy consumption by 11.6%.

automatic code generationCUDAenergy efficiency

This study addresses hardware obsolescence caused by software bloat and GPU performance bottlenecks in resource-constrained virtual machines. We propose deploying the Uxn virtual machine on integrated graphics and introduce an OpenMP-style parallel API based on the Uxntal language. By exploiting data parallelism, this approach enables the efficient execution of frugal computing workloads on general-purpose GPUs. Experimental evaluations demonstrate a 19× speedup on stencil benchmarks and a 7× increase in Bunnymark frame rates. These results effectively validate the feasibility and significant advantages of leveraging GPU acceleration for frugal computing within resource-limited environments, offering a viable solution to extend hardware lifecycle while maintaining computational efficiency.

GPU implementationResource-constrained VMSoftware bloat

This study addresses the low utilization of modern GPU computing resources by systematically evaluating the performance, energy efficiency, and resource isolation characteristics of NVIDIA’s Multi-Process Service (MPS) and Multi-Instance GPU (MIG) technologies under concurrent application workloads. The experiments reveal a critical trade-off between MPS’s scheduling flexibility and MIG’s hardware-level isolation: MPS can improve performance by up to 30% and reduce energy consumption by approximately 20% in the absence of memory contention, yet suffers a 30% performance degradation under contention; MIG effectively mitigates resource contention but is constrained by its rigid configuration options and higher overhead. These findings provide empirical foundations for optimizing GPU co-execution strategies driven by application-specific workload characteristics.

co-executionGPU underutilizationperformance isolation