gpu architecture

Designs, implements, or analyzes the microarchitecture and system-level organization of graphics processing units, including compute cores/streaming multiprocessors, SIMT/SIMD execution model, instruction scheduling, memory hierarchy and caches, interconnects, and hardware support for concurrency, power/performance tradeoffs, and accelerator integration or virtualization.

gpuarchitecture

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.33
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$215K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the lack of a unified instruction set architecture across GPU vendors, which hinders efficient cross-platform portability of parallel programs. Through a systematic analysis of instruction sets from sixteen microarchitectures spanning four major vendors, the work identifies ten cross-platform computational primitives, six dialect-like variations, and six fundamental architectural divergences. Leveraging these insights, it proposes the first vendor-agnostic abstract execution model for GPUs. Validated against official documentation, patents, reverse-engineered data, and cross-platform benchmarks, the model demonstrates strong performance on architecturally disparate hardware—specifically NVIDIA T4 and Apple M1—matching or exceeding native performance in five out of six benchmark suites, with only parallel reduction lagging at 62.5% efficiency, thereby underscoring the critical role of the shuffle primitive.

computational primitivescross-vendorGPU ISA

Analyzing Modern NVIDIA GPU cores

Mar 26, 2025
RH
Rodrigo Huerta
🏛️ Universitat Politècnica de Catalunya

Contemporary NVIDIA GPU microarchitectural research lags significantly, often relying on designs over fifteen years old. Method: This paper presents the first systematic reverse-engineering study of RTX-class GPU cores, uncovering their instruction scheduling policies, register file and cache hierarchies, memory pipeline characteristics, and hardware-software co-execution mechanisms. It proposes a stream-buffer-based instruction prefetcher and empirically demonstrates that software-managed dependency tracking outperforms traditional hardware scoreboard approaches. Contribution/Results: We develop a high-fidelity instruction-level simulator incorporating detailed register file caching and read-port modeling. Evaluated on the RTX A6000, it achieves a mean absolute percentage error (MAPE) of 13.98%, representing an 18.24% improvement over state-of-the-art simulators. Moreover, the model exhibits cross-generational generalizability, successfully transferring to the Turing architecture.

Analyze instruction prefetcher and register file impactImprove simulation accuracy for HPC workloadsReverse engineer modern NVIDIA GPU cores' design

This study addresses the absence of open standards for CPU pipeline visualization tools and the difficulty in localizing performance bottlenecks. To this end, it proposes an open-source event stream format alongside Catscan, an interactive viewer. Methodologically, this work introduces a structured event stream based on transactional relationships, integrating typed event modeling, persistent highlighting techniques, and domain-specific search algorithms to enable microarchitectural trace analysis from symptoms down to individual instructions. Furthermore, it supports resource-oriented views synchronized with comparative trace alignment. By successfully reproducing industry-grade debugging workflows, this project provides the community with production-validated microarchitectural visualization infrastructure.

CPU performance simulationmicroarchitecture debuggingopen-source tooling

Concurrent Scheduling of High-Level Parallel Programs on Multi-GPU Systems

Mar 13, 2025
FK
Fabian Knorr
🏛️ University of Innsbruck

SYCL programs on multi-GPU clusters suffer from high scheduling latency and substantial critical-path overhead due to implicit memory allocation, cache-coherence operations, and dependency analysis. Method: We propose the Instruction Graph—a novel intermediate representation that fully decouples scheduling from execution. Our approach integrates speculative scheduling, adaptive virtual-buffer memory allocation, and tight integration with the Celerity runtime, enabling fully concurrent scheduling of memory management, data transfers, MPI communication, and kernel launches while moving all scheduling analysis off the critical execution path. Contribution/Results: Evaluated on a production-scale 128-GPU cluster, our method achieves excellent strong scaling, drastically reduces multi-application scheduling latency, and drives critical-path overhead nearly to zero—thereby overcoming fundamental limitations of conventional static and blocking schedulers.

Enhancing memory allocation and concurrency in SYCL programs on accelerator clusters.Optimizing scheduling for high-level parallel programs on multi-GPU systems.Reducing delays in distributed-memory applications through graph-based representations.

FPGA or GPU? Analyzing Comparative Research for Application-Specific Guidance

Mar 22, 2025
AA
Arnab A Purkayastha
🏛️ Western New England University | The Citadel

Existing FPGA–GPU comparative studies predominantly focus on raw performance metrics and lack domain-specific guidance for accelerator selection. Method: This paper proposes an application-oriented, fine-grained comparative framework that systematically synthesizes over 100 studies, conducting cross-domain (e.g., AI, HPC, network processing, scientific computing) classification and cross-evaluation along three dimensions: performance, energy efficiency, and programmability. Contribution/Results: The study innovatively establishes the first empirically grounded applicability boundaries for FPGAs and GPUs: FPGAs excel in low-latency, high-throughput customized pipelines and energy-constrained scenarios; GPUs are superior for massively parallel, computation-intensive workloads with stable algorithms. The resulting actionable decision-making guide enables researchers and engineers to select hardware accelerators based on domain-specific requirements, thereby bridging the gap between architectural characteristics and real-world application needs.

Analyzing FPGA vs GPU performance for application-specific guidanceBridging research gap on ideal accelerators for domain-specific applicationsProviding actionable recommendations for hardware accelerator selection

Latest Papers

What's happening recently
View more

This study addresses the low utilization of modern GPU computing resources by systematically evaluating the performance, energy efficiency, and resource isolation characteristics of NVIDIA’s Multi-Process Service (MPS) and Multi-Instance GPU (MIG) technologies under concurrent application workloads. The experiments reveal a critical trade-off between MPS’s scheduling flexibility and MIG’s hardware-level isolation: MPS can improve performance by up to 30% and reduce energy consumption by approximately 20% in the absence of memory contention, yet suffers a 30% performance degradation under contention; MIG effectively mitigates resource contention but is constrained by its rigid configuration options and higher overhead. These findings provide empirical foundations for optimizing GPU co-execution strategies driven by application-specific workload characteristics.

co-executionGPU underutilizationperformance isolation

This work systematically investigates the trade-offs between performance and programmability in CPU-GPU cooperative scheduling across discrete and unified memory architectures, with a focus on sparse conjugate gradient computations. Evaluations are conducted on both the NVIDIA GH200 Superchip—a platform featuring a unified memory architecture—and the discrete H100 PCIe system, comparing three memory management paradigms: explicit data copies, managed memory, and mapped memory. The study reveals that the GH200’s fused architecture substantially enhances the practicality of managed memory, enabling diverse hybrid task-partitioning strategies to achieve both high performance and programming simplicity. These findings underscore the significant impact of underlying memory architecture on the efficacy of cooperative scheduling approaches.

coschedulingCPU-GPUintegrated GPU

This work addresses the high energy consumption of traditional GPUs under the SIMT programming model, which stems from frequent register file accesses and complex control logic. The authors propose replacing the SIMD backend with a statically scheduled coarse-grained reconfigurable array (CGRA) that pipelines active threads and enables direct data transfer among processing elements, drastically reducing intermediate value accesses to registers. A novel p-graph program representation is introduced to decouple dynamic dependency edges, and in conjunction with double-buffered configuration memory, compile-time graph unrolling, and a temporal memory coalescing unit (TMCU), the design efficiently supports dynamic behavior while retaining static scheduling. Experimental results on the Rodinia benchmark suite show an average 68% reduction in register file accesses, 1.77–1.90× improvement in dynamic energy efficiency, 42.0%–45.9% lower power consumption, and performance comparable to that of NVIDIA Turing GPUs.

coarse-grained reconfigurable arraysenergy overheadregister file accesses

Current GPU programming models lack expressiveness for chiplet-level locality and synchronization, leading to redundant memory accesses and poor cache utilization when executing memory-intensive workloads such as large language model (LLM) inference on multi-chiplet GPUs. This work proposes Fleet, the first multi-level task programming model that explicitly exposes the chiplet hierarchy. Fleet introduces a chiplet-task abstraction that binds computation and data to specific chiplets and integrates persistent kernels, cooperative weight tiling, and per-chiplet scheduling to enable L2 cache reuse and efficient coordinated execution. Evaluated on an AMD MI350 running Qwen3-8B, Fleet reduces decoding latency by 1.3–1.5× for small batches and cuts HBM traffic by up to 37% under large batches, significantly improving L2 hit rates and achieving overall speedups of 1.27–1.30×.

cache utilizationchipletGPU programming model

Hot Scholars

XH

Xiaolin Huang

Professor, Shanghai Jiao Tong University
machine learningkernel methoddeep neural network trainingpiecewise linear model
MS

Maximilian Schiffer

School of Management & Munich Data Science Institute, Technical University of Munich
operations researchsmart mobilitytransportationindustrial engineering
PH

Pedro Hermosilla

TU Wien
3D Machine Learning3D Computer VisionPoint Cloud ProcessingComputer Graphics
JL

Jingwen Leng

Professor, Shanghai Jiao Tong University
Computer Architecture
XL

Xinhao Luo

Ph.D. student, Shanghai Jiao Tong University
High Performance ComputingML Compiler