Score
Designs and implements device-specific backend software that interfaces with low-level graphics/compute APIs such as Metal to expose and manage hardware capabilities. Work includes command encoding and submission, resource and memory management, synchronization primitives, pipeline and shader binding, integration with host runtimes, and performance debugging and optimization to ensure correct, efficient execution on the target device.
Existing NPU low-level programming faces a trade-off between high development costs of low-level tools and excessive abstraction in high-level tools, which forfeits critical optimizations. Method: This paper introduces IRON—a next-generation near-metal programming interface—featuring a unified programming abstraction that balances efficiency, expressiveness, and extensibility; a novel modular, pluggable placement-and-tiling toolchain architecture; and integrated techniques including domain-specific IR interfaces, Halstead-complexity-driven code simplification, NPU hardware-coordinated modeling, and dataflow graph layout optimization. Contribution/Results: IRON reduces average code size by 26% and significantly lowers Halstead complexity while maintaining full backward compatibility with prior IRON functionality. It supports multiple mainstream AI acceleration paradigms and substantially improves NPU performance engineering productivity.
Heterogeneous GPUs exhibit severe binary incompatibility due to divergent instruction sets, execution models, and driver stacks. To address this, we present the first cross-vendor binary-compatible system supporting NVIDIA, AMD, Intel, and Tenstorrent GPUs. Our approach comprises three key components: (1) an architecture-agnostic intermediate representation (IR) that unifies SIMT and MIMD execution semantics; (2) a dynamic binary translation runtime coupled with state-aware serialization for real-time migration; and (3) a unified abstraction layer ensuring consistent cross-platform semantics for threads, memory, and synchronization. Crucially, our system enables direct binary migration across all four GPU architectures without source or binary modification. Evaluation shows migration overhead under 8% and average performance degradation below 12%, effectively breaking down long-standing hardware-software co-design barriers across GPU vendors.
This work addresses the lack of transparency in NVIDIA’s closed-source user-space driver, which obscures the translation of CUDA API calls into hardware commands and impedes understanding of GPU behavior and performance attribution. The authors propose a novel approach that requires no modification to the proprietary driver, instead leveraging an open-source kernel driver, memory-mapped path instrumentation, and hardware watchpoints on the GPU’s doorbell registers to capture and reconstruct the complete low-level command stream with unprecedented accuracy. This methodology reveals the true DMA patterns and performance characteristics of CUDA data transfers and demonstrates that the low overhead of CUDA Graphs stems from their streamlined and efficient command submission mechanism. By significantly enhancing the interpretability of GPU runtime behavior, this approach establishes a new paradigm for middleware analysis and hardware-software co-design.
This study addresses the lack of a unified instruction set architecture across GPU vendors, which hinders efficient cross-platform portability of parallel programs. Through a systematic analysis of instruction sets from sixteen microarchitectures spanning four major vendors, the work identifies ten cross-platform computational primitives, six dialect-like variations, and six fundamental architectural divergences. Leveraging these insights, it proposes the first vendor-agnostic abstract execution model for GPUs. Validated against official documentation, patents, reverse-engineered data, and cross-platform benchmarks, the model demonstrates strong performance on architecturally disparate hardware—specifically NVIDIA T4 and Apple M1—matching or exceeding native performance in five out of six benchmark suites, with only parallel reduction lagging at 62.5% efficiency, thereby underscoring the critical role of the shuffle primitive.
Assessing performance portability across GPU programming models on heterogeneous NVIDIA and AMD hardware remains challenging due to fragmented benchmarks and irreproducible evaluation methodologies. Method: We conduct a systematic, cross-platform evaluation of seven programming models—CUDA, HIP, Kokkos, RAJA, OpenMP, OpenACC, and SYCL—using five interdisciplinary proxy applications. Leveraging a Spack-based automation framework, we ensure fully reproducible build, deployment, and benchmarking across real multi-vendor GPU systems. Contribution/Results: Our empirical study quantifies both performance consistency and migration overhead for each model. HIP and SYCL achieve superior performance on AMD GPUs; CUDA remains dominant on NVIDIA hardware; Kokkos and RAJA deliver balanced portability with moderate performance; OpenMP and OpenACC exhibit significant cross-platform performance degradation. This work provides the first unified, vendor-agnostic assessment of GPU programming models’ performance portability and establishes a rigorous, reproducible methodology to guide architecture selection for high-performance scientific software.
Selecting optimal hardware for deploying small language models (SLMs) in resource-constrained edge computing environments remains challenging due to the lack of systematic, cross-architecture performance and efficiency comparisons. Method: This work conducts the first unified, empirical evaluation—within a consistent experimental framework—of Intel/ARM CPUs, NVIDIA GPUs, and the RaiderChip NPU across mainstream SLMs. We introduce bandwidth-normalized analysis and adopt holistic metrics—including Energy-Delay Product (EDP)—to jointly quantify inference throughput, latency, and energy efficiency. Results: Dedicated NPUs achieve substantially higher throughput and 1–2 orders-of-magnitude lower EDP than general-purpose CPUs. While low-power ARM CPUs exhibit limited raw performance, they demonstrate competitive inference-per-watt efficiency. Our reproducible benchmarking methodology and empirical findings provide actionable, evidence-based guidance for hardware selection in edge-deployed SLM applications.
Modern GPU systems suffer from inflexible static resource management, hindering efficient adaptation to diverse workloads: user-space runtimes lack cross-tenant visibility and hardware control, while kernel-level modifications introduce security vulnerabilities and maintenance overhead. This paper introduces the first eBPF-based policy runtime for GPUs, abstracting GPU drivers and hardware as a programmable OS subsystem. Our key contributions are: (1) a lightweight device-side eBPF virtual machine enabling safe execution of verified policies within the GPU kernel; and (2) a secure, driver-level hooking mechanism that jointly ensures programmability, fine-grained hardware control, and multi-tenant observability. Evaluated on inference, training, and vector search workloads, our approach achieves up to 4.8× higher throughput and 2× lower tail latency. Policy deployment requires zero application modification and zero driver restarts, with runtime overhead under 3%.
This study addresses the low utilization of modern GPU computing resources by systematically evaluating the performance, energy efficiency, and resource isolation characteristics of NVIDIA’s Multi-Process Service (MPS) and Multi-Instance GPU (MIG) technologies under concurrent application workloads. The experiments reveal a critical trade-off between MPS’s scheduling flexibility and MIG’s hardware-level isolation: MPS can improve performance by up to 30% and reduce energy consumption by approximately 20% in the absence of memory contention, yet suffers a 30% performance degradation under contention; MIG effectively mitigates resource contention but is constrained by its rigid configuration options and higher overhead. These findings provide empirical foundations for optimizing GPU co-execution strategies driven by application-specific workload characteristics.
This study addresses hardware obsolescence caused by software bloat and GPU performance bottlenecks in resource-constrained virtual machines. We propose deploying the Uxn virtual machine on integrated graphics and introduce an OpenMP-style parallel API based on the Uxntal language. By exploiting data parallelism, this approach enables the efficient execution of frugal computing workloads on general-purpose GPUs. Experimental evaluations demonstrate a 19× speedup on stencil benchmarks and a 7× increase in Bunnymark frame rates. These results effectively validate the feasibility and significant advantages of leveraging GPU acceleration for frugal computing within resource-limited environments, offering a viable solution to extend hardware lifecycle while maintaining computational efficiency.
This work presents the first empirical investigation into the execution mechanism of fp8 matmul2d operations in Apple Metal 4.1 on the M4 Max GPU, revealing that they are implemented via software emulation on shader cores rather than dedicated hardware acceleration, utilizing an undocumented 8×8 tensor tiling layout. Through a meticulously designed microbenchmarking framework incorporating checksum-based validation, provenance tracking, throughput ceiling analysis, comparison against simdgroup_matrix primitives, and power attribution, we demonstrate that fp8 achieves only 0.94× the performance of fp16, confirming its primary benefit lies in memory footprint reduction. Leveraging these insights, we develop a hand-optimized fused kernel that attains performance gains of 6.5–12.9% in cache-resident scenarios.