device backend implementation

Designs and implements device-specific backend software that interfaces with low-level graphics/compute APIs such as Metal to expose and manage hardware capabilities. Work includes command encoding and submission, resource and memory management, synchronization primitives, pipeline and shader binding, integration with host runtimes, and performance debugging and optimization to ensure correct, efficient execution on the target device.

devicebackendimplementation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.03
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Efficiency, Expressivity, and Extensibility in a Close-to-Metal NPU Programming Interface

Apr 25, 2025
EH
Erika Hunhoff
🏛️ University of Colorado | AMD

Existing NPU low-level programming faces a trade-off between high development costs of low-level tools and excessive abstraction in high-level tools, which forfeits critical optimizations. Method: This paper introduces IRON—a next-generation near-metal programming interface—featuring a unified programming abstraction that balances efficiency, expressiveness, and extensibility; a novel modular, pluggable placement-and-tiling toolchain architecture; and integrated techniques including domain-specific IR interfaces, Halstead-complexity-driven code simplification, NPU hardware-coordinated modeling, and dataflow graph layout optimization. Contribution/Results: IRON reduces average code size by 26% and significantly lowers Halstead complexity while maintaining full backward compatibility with prior IRON functionality. It supports multiple mainstream AI acceleration paradigms and substantially improves NPU performance engineering productivity.

Balancing expressivity and optimization in close-to-metal toolkitsEnhancing NPU programming efficiency with reduced code complexityExtending interface features for flexible data and placement control

Heterogeneous GPUs exhibit severe binary incompatibility due to divergent instruction sets, execution models, and driver stacks. To address this, we present the first cross-vendor binary-compatible system supporting NVIDIA, AMD, Intel, and Tenstorrent GPUs. Our approach comprises three key components: (1) an architecture-agnostic intermediate representation (IR) that unifies SIMT and MIMD execution semantics; (2) a dynamic binary translation runtime coupled with state-aware serialization for real-time migration; and (3) a unified abstraction layer ensuring consistent cross-platform semantics for threads, memory, and synchronization. Crucially, our system enables direct binary migration across all four GPU architectures without source or binary modification. Evaluation shows migration overhead under 8% and average performance degradation below 12%, effectively breaking down long-standing hardware-software co-design barriers across GPU vendors.

Bridging execution model differences for seamless GPU migrationDynamic translation of GPU intermediate representation to native codeEnabling binary compatibility across diverse GPU vendors

This work addresses the lack of transparency in NVIDIA’s closed-source user-space driver, which obscures the translation of CUDA API calls into hardware commands and impedes understanding of GPU behavior and performance attribution. The authors propose a novel approach that requires no modification to the proprietary driver, instead leveraging an open-source kernel driver, memory-mapped path instrumentation, and hardware watchpoints on the GPU’s doorbell registers to capture and reconstruct the complete low-level command stream with unprecedented accuracy. This methodology reveals the true DMA patterns and performance characteristics of CUDA data transfers and demonstrates that the low overhead of CUDA Graphs stems from their streamlined and efficient command submission mechanism. By significantly enhancing the interpretability of GPU runtime behavior, this approach establishes a new paradigm for middleware analysis and hardware-software co-design.

closed-source drivercommand streamCUDA

This study addresses the lack of a unified instruction set architecture across GPU vendors, which hinders efficient cross-platform portability of parallel programs. Through a systematic analysis of instruction sets from sixteen microarchitectures spanning four major vendors, the work identifies ten cross-platform computational primitives, six dialect-like variations, and six fundamental architectural divergences. Leveraging these insights, it proposes the first vendor-agnostic abstract execution model for GPUs. Validated against official documentation, patents, reverse-engineered data, and cross-platform benchmarks, the model demonstrates strong performance on architecturally disparate hardware—specifically NVIDIA T4 and Apple M1—matching or exceeding native performance in five out of six benchmark suites, with only parallel reduction lagging at 62.5% efficiency, thereby underscoring the critical role of the shuffle primitive.

computational primitivescross-vendorGPU ISA

Taking GPU Programming Models to Task for Performance Portability

Feb 14, 2024
JH
Joshua H. Davis
🏛️ University of Maryland | Lawrence Livermore National Laboratory

Assessing performance portability across GPU programming models on heterogeneous NVIDIA and AMD hardware remains challenging due to fragmented benchmarks and irreproducible evaluation methodologies. Method: We conduct a systematic, cross-platform evaluation of seven programming models—CUDA, HIP, Kokkos, RAJA, OpenMP, OpenACC, and SYCL—using five interdisciplinary proxy applications. Leveraging a Spack-based automation framework, we ensure fully reproducible build, deployment, and benchmarking across real multi-vendor GPU systems. Contribution/Results: Our empirical study quantifies both performance consistency and migration overhead for each model. HIP and SYCL achieve superior performance on AMD GPUs; CUDA remains dominant on NVIDIA hardware; Kokkos and RAJA deliver balanced portability with moderate performance; OpenMP and OpenACC exhibit significant cross-platform performance degradation. This work provides the first unified, vendor-agnostic assessment of GPU programming models’ performance portability and establishes a rigorous, reproducible methodology to guide architecture selection for high-performance scientific software.

Analyzing underperformance causes and providing optimizationsAssessing consistency on NVIDIA and AMD GPU architecturesEvaluating performance portability across GPU programming models

Latest Papers

What's happening recently
View more

Selecting optimal hardware for deploying small language models (SLMs) in resource-constrained edge computing environments remains challenging due to the lack of systematic, cross-architecture performance and efficiency comparisons. Method: This work conducts the first unified, empirical evaluation—within a consistent experimental framework—of Intel/ARM CPUs, NVIDIA GPUs, and the RaiderChip NPU across mainstream SLMs. We introduce bandwidth-normalized analysis and adopt holistic metrics—including Energy-Delay Product (EDP)—to jointly quantify inference throughput, latency, and energy efficiency. Results: Dedicated NPUs achieve substantially higher throughput and 1–2 orders-of-magnitude lower EDP than general-purpose CPUs. While low-power ARM CPUs exhibit limited raw performance, they demonstrate competitive inference-per-watt efficiency. Our reproducible benchmarking methodology and empirical findings provide actionable, evidence-based guidance for hardware selection in edge-deployed SLM applications.

Analyzes energy efficiency and speed trade-offs for edge AI deployment.Compares hardware backends for SLM inference under strict resource constraints.Evaluates CPU, GPU, NPU performance for Small Language Models on edge devices.

gpu_ext: Extensible OS Policies for GPUs via eBPF

Dec 14, 2025
YZ
Yusheng Zheng
🏛️ UC Santa Cruz | Alibaba Group | University of Washington | University of Connecticut | ShanghaiTech University | Virginia Tech

Modern GPU systems suffer from inflexible static resource management, hindering efficient adaptation to diverse workloads: user-space runtimes lack cross-tenant visibility and hardware control, while kernel-level modifications introduce security vulnerabilities and maintenance overhead. This paper introduces the first eBPF-based policy runtime for GPUs, abstracting GPU drivers and hardware as a programmable OS subsystem. Our key contributions are: (1) a lightweight device-side eBPF virtual machine enabling safe execution of verified policies within the GPU kernel; and (2) a secure, driver-level hooking mechanism that jointly ensures programmability, fine-grained hardware control, and multi-tenant observability. Evaluated on inference, training, and vector search workloads, our approach achieves up to 4.8× higher throughput and 2× lower tail latency. Policy deployment requires zero application modification and zero driver restarts, with runtime overhead under 3%.

Enabling safe and efficient device-side policy execution via eBPFExtensible OS policies for GPU resource managementOvercoming limitations of user-space runtimes and kernel modifications

This study addresses the low utilization of modern GPU computing resources by systematically evaluating the performance, energy efficiency, and resource isolation characteristics of NVIDIA’s Multi-Process Service (MPS) and Multi-Instance GPU (MIG) technologies under concurrent application workloads. The experiments reveal a critical trade-off between MPS’s scheduling flexibility and MIG’s hardware-level isolation: MPS can improve performance by up to 30% and reduce energy consumption by approximately 20% in the absence of memory contention, yet suffers a 30% performance degradation under contention; MIG effectively mitigates resource contention but is constrained by its rigid configuration options and higher overhead. These findings provide empirical foundations for optimizing GPU co-execution strategies driven by application-specific workload characteristics.

co-executionGPU underutilizationperformance isolation

This study addresses hardware obsolescence caused by software bloat and GPU performance bottlenecks in resource-constrained virtual machines. We propose deploying the Uxn virtual machine on integrated graphics and introduce an OpenMP-style parallel API based on the Uxntal language. By exploiting data parallelism, this approach enables the efficient execution of frugal computing workloads on general-purpose GPUs. Experimental evaluations demonstrate a 19× speedup on stencil benchmarks and a 7× increase in Bunnymark frame rates. These results effectively validate the feasibility and significant advantages of leveraging GPU acceleration for frugal computing within resource-limited environments, offering a viable solution to extend hardware lifecycle while maintaining computational efficiency.

GPU implementationResource-constrained VMSoftware bloat

This work presents the first empirical investigation into the execution mechanism of fp8 matmul2d operations in Apple Metal 4.1 on the M4 Max GPU, revealing that they are implemented via software emulation on shader cores rather than dedicated hardware acceleration, utilizing an undocumented 8×8 tensor tiling layout. Through a meticulously designed microbenchmarking framework incorporating checksum-based validation, provenance tracking, throughput ceiling analysis, comparison against simdgroup_matrix primitives, and power attribution, we demonstrate that fp8 achieves only 0.94× the performance of fp16, confirming its primary benefit lies in memory footprint reduction. Leveraging these insights, we develop a hand-optimized fused kernel that attains performance gains of 6.5–12.9% in cache-resident scenarios.

Apple M4 GPUhardware accelerationmatrix multiplication