Score
Designs, builds, and evaluates systems and software that use NVIDIA GPUs, including writing and optimizing CUDA/OpenCL kernels, managing drivers and libraries, and integrating multi‑GPU configurations. Involves profiling and benchmarking, tuning memory and compute concurrency, and diagnosing hardware and software performance, scalability, and power/thermal behavior for GPU‑accelerated applications.
This work addresses the lack of transparency in NVIDIA’s closed-source user-space driver, which obscures the translation of CUDA API calls into hardware commands and impedes understanding of GPU behavior and performance attribution. The authors propose a novel approach that requires no modification to the proprietary driver, instead leveraging an open-source kernel driver, memory-mapped path instrumentation, and hardware watchpoints on the GPU’s doorbell registers to capture and reconstruct the complete low-level command stream with unprecedented accuracy. This methodology reveals the true DMA patterns and performance characteristics of CUDA data transfers and demonstrates that the low overhead of CUDA Graphs stems from their streamlined and efficient command submission mechanism. By significantly enhancing the interpretability of GPU runtime behavior, this approach establishes a new paradigm for middleware analysis and hardware-software co-design.
This work addresses the critical gap in existing performance analysis tools, which largely lack support for virtual GPUs (vGPUs) and thus struggle to diagnose performance bottlenecks in vGPU-accelerated applications. To tackle this limitation, the authors design and implement a novel vGPU performance profiling tool tailored for Intel GVT-g. The tool introduces, for the first time, a fine-grained software tracing mechanism coupled with runtime data collection to generate multidimensional performance metrics. These metrics are then presented through synchronized multi-view visualizations that capture comprehensive vGPU behavioral characteristics. By enabling detailed insight into vGPU execution dynamics, this approach significantly enhances observability and debugging efficiency in GPU-virtualized environments, effectively filling a longstanding void in vGPU performance analysis tooling.
GPU topology information—such as compute/memory hierarchy, cache sizes, interconnect bandwidths, and physical layout—is severely fragmented, incomplete, and vendor-locked, hindering performance modeling and optimization in HPC and AI systems. To address this, we propose the first cross-vendor (NVIDIA/AMD), open-source framework for automatic GPU hardware topology discovery. Our approach integrates CUDA/HIP APIs with over 50 customized microbenchmarks and employs statistical validation—including Kolmogorov–Smirnov tests—to infer non-programmable hardware attributes with high confidence. We validate the framework across ten mainstream GPUs, demonstrating broad compatibility and measurement accuracy. Furthermore, we integrate it into three critical workflows: analytical performance modeling, bottleneck analysis, and dynamic resource partitioning. This integration significantly enhances system-level hardware awareness and improves resource utilization efficiency.
This study addresses the GPU computational efficiency bottleneck in deep and machine learning. Methodologically, it proposes a task-aware GPU parallel architecture adaptation framework that systematically integrates CUDA stream-based concurrency, dynamic parallelism, and heterogeneous hardware (FPGA/TPU/ASIC) co-selection—implemented via deep integration into PyTorch, TensorFlow, and XGBoost. Its core contribution lies in establishing a transferable GPU optimization methodology, transcending model- or library-specific tuning. Experimental evaluation demonstrates 3–8× speedup across representative training and inference workloads. Furthermore, the authors open-source a modular, well-documented GPU optimization practice guide, substantially lowering the barrier to entry for AI practitioners seeking parallelization optimizations.
Assessing performance portability across GPU programming models on heterogeneous NVIDIA and AMD hardware remains challenging due to fragmented benchmarks and irreproducible evaluation methodologies. Method: We conduct a systematic, cross-platform evaluation of seven programming models—CUDA, HIP, Kokkos, RAJA, OpenMP, OpenACC, and SYCL—using five interdisciplinary proxy applications. Leveraging a Spack-based automation framework, we ensure fully reproducible build, deployment, and benchmarking across real multi-vendor GPU systems. Contribution/Results: Our empirical study quantifies both performance consistency and migration overhead for each model. HIP and SYCL achieve superior performance on AMD GPUs; CUDA remains dominant on NVIDIA hardware; Kokkos and RAJA deliver balanced portability with moderate performance; OpenMP and OpenACC exhibit significant cross-platform performance degradation. This work provides the first unified, vendor-agnostic assessment of GPU programming models’ performance portability and establishes a rigorous, reproducible methodology to guide architecture selection for high-performance scientific software.
This work addresses three key challenges in real-time GPU performance monitoring at planetary scale: user privacy leakage, runtime overhead, and difficulty in data attribution across massive deployments. We propose the first end-to-end privacy-preserving GPU performance profiling architecture that operates at planetary scale with zero performance overhead. Methodologically, we design a lightweight kernel-level monitoring agent—built upon NSYS/NCU—that integrates differential privacy, secure aggregation, and distributed sampling to enable application-agnostic, anonymized kernel-level behavior attribution and resource accounting. Evaluated on a simulated 100,000-GPU cluster, our system achieves complete telemetry collection while enabling precise application-level attribution for realistic deep learning workloads (e.g., TorchBench) with no third-party privacy leakage. The architecture delivers trustworthy, scalable hardware performance insights for chip design and systems optimization.
Existing LLM-driven approaches for GPU kernel optimization are largely confined to machine learning scenarios, lacking cross-domain generality and a systematic evaluation benchmark. To address this gap, this work proposes CUDAMaster—a hardware-aware, multi-agent automated optimization framework that integrates performance profiling with an automated compilation toolchain to generate CUDA kernels across diverse workloads under FP32 and BF16 precision. Furthermore, we introduce MSKernelBench, the first comprehensive benchmark encompassing algebraic operations, LLM operators, sparse matrix computations, and scientific computing. Experimental results demonstrate that CUDAMaster consistently outperforms Astra by an average of approximately 35% across most operators, with certain kernels matching or even surpassing the performance of the proprietary cuBLAS library.
This work proposes KernelPro, a closed-loop multi-agent system designed to automatically generate high-performance and energy-efficient GPU kernel code. By integrating large language models with hardware micro-benchmarking tools, KernelPro employs semantic feedback operators, a two-tier tool-calling architecture, a domain-adapted Monte Carlo Tree Search (MCTS) strategy, and direct CuTe source-code generation to jointly optimize for both performance and energy efficiency. Evaluated on KernelBench, KernelPro achieves up to a 5.30× speedup over baseline implementations. Furthermore, when applied to expert-optimized Mixture-of-Experts (MoE) kernels, it outperforms hand-tuned Triton kernels by 1.23× in performance while reducing measured energy consumption by 11.6%.
This study addresses hardware obsolescence caused by software bloat and GPU performance bottlenecks in resource-constrained virtual machines. We propose deploying the Uxn virtual machine on integrated graphics and introduce an OpenMP-style parallel API based on the Uxntal language. By exploiting data parallelism, this approach enables the efficient execution of frugal computing workloads on general-purpose GPUs. Experimental evaluations demonstrate a 19× speedup on stencil benchmarks and a 7× increase in Bunnymark frame rates. These results effectively validate the feasibility and significant advantages of leveraging GPU acceleration for frugal computing within resource-limited environments, offering a viable solution to extend hardware lifecycle while maintaining computational efficiency.
This study addresses the low utilization of modern GPU computing resources by systematically evaluating the performance, energy efficiency, and resource isolation characteristics of NVIDIA’s Multi-Process Service (MPS) and Multi-Instance GPU (MIG) technologies under concurrent application workloads. The experiments reveal a critical trade-off between MPS’s scheduling flexibility and MIG’s hardware-level isolation: MPS can improve performance by up to 30% and reduce energy consumption by approximately 20% in the absence of memory contention, yet suffers a 30% performance degradation under contention; MIG effectively mitigates resource contention but is constrained by its rigid configuration options and higher overhead. These findings provide empirical foundations for optimizing GPU co-execution strategies driven by application-specific workload characteristics.