Score
Designs, implements, or analyzes the microarchitecture and system-level organization of graphics processing units, including compute cores/streaming multiprocessors, SIMT/SIMD execution model, instruction scheduling, memory hierarchy and caches, interconnects, and hardware support for concurrency, power/performance tradeoffs, and accelerator integration or virtualization.
This study addresses the lack of a unified instruction set architecture across GPU vendors, which hinders efficient cross-platform portability of parallel programs. Through a systematic analysis of instruction sets from sixteen microarchitectures spanning four major vendors, the work identifies ten cross-platform computational primitives, six dialect-like variations, and six fundamental architectural divergences. Leveraging these insights, it proposes the first vendor-agnostic abstract execution model for GPUs. Validated against official documentation, patents, reverse-engineered data, and cross-platform benchmarks, the model demonstrates strong performance on architecturally disparate hardware—specifically NVIDIA T4 and Apple M1—matching or exceeding native performance in five out of six benchmark suites, with only parallel reduction lagging at 62.5% efficiency, thereby underscoring the critical role of the shuffle primitive.
Contemporary NVIDIA GPU microarchitectural research lags significantly, often relying on designs over fifteen years old. Method: This paper presents the first systematic reverse-engineering study of RTX-class GPU cores, uncovering their instruction scheduling policies, register file and cache hierarchies, memory pipeline characteristics, and hardware-software co-execution mechanisms. It proposes a stream-buffer-based instruction prefetcher and empirically demonstrates that software-managed dependency tracking outperforms traditional hardware scoreboard approaches. Contribution/Results: We develop a high-fidelity instruction-level simulator incorporating detailed register file caching and read-port modeling. Evaluated on the RTX A6000, it achieves a mean absolute percentage error (MAPE) of 13.98%, representing an 18.24% improvement over state-of-the-art simulators. Moreover, the model exhibits cross-generational generalizability, successfully transferring to the Turing architecture.
This study addresses the absence of open standards for CPU pipeline visualization tools and the difficulty in localizing performance bottlenecks. To this end, it proposes an open-source event stream format alongside Catscan, an interactive viewer. Methodologically, this work introduces a structured event stream based on transactional relationships, integrating typed event modeling, persistent highlighting techniques, and domain-specific search algorithms to enable microarchitectural trace analysis from symptoms down to individual instructions. Furthermore, it supports resource-oriented views synchronized with comparative trace alignment. By successfully reproducing industry-grade debugging workflows, this project provides the community with production-validated microarchitectural visualization infrastructure.
SYCL programs on multi-GPU clusters suffer from high scheduling latency and substantial critical-path overhead due to implicit memory allocation, cache-coherence operations, and dependency analysis. Method: We propose the Instruction Graph—a novel intermediate representation that fully decouples scheduling from execution. Our approach integrates speculative scheduling, adaptive virtual-buffer memory allocation, and tight integration with the Celerity runtime, enabling fully concurrent scheduling of memory management, data transfers, MPI communication, and kernel launches while moving all scheduling analysis off the critical execution path. Contribution/Results: Evaluated on a production-scale 128-GPU cluster, our method achieves excellent strong scaling, drastically reduces multi-application scheduling latency, and drives critical-path overhead nearly to zero—thereby overcoming fundamental limitations of conventional static and blocking schedulers.
Existing FPGA–GPU comparative studies predominantly focus on raw performance metrics and lack domain-specific guidance for accelerator selection. Method: This paper proposes an application-oriented, fine-grained comparative framework that systematically synthesizes over 100 studies, conducting cross-domain (e.g., AI, HPC, network processing, scientific computing) classification and cross-evaluation along three dimensions: performance, energy efficiency, and programmability. Contribution/Results: The study innovatively establishes the first empirically grounded applicability boundaries for FPGAs and GPUs: FPGAs excel in low-latency, high-throughput customized pipelines and energy-constrained scenarios; GPUs are superior for massively parallel, computation-intensive workloads with stable algorithms. The resulting actionable decision-making guide enables researchers and engineers to select hardware accelerators based on domain-specific requirements, thereby bridging the gap between architectural characteristics and real-world application needs.
为解决GPU编程中性能优化与安全性问题,提出Exo-GPU语言,通过将并行性和同步性作为顺序代码注解处理,保证功能等效同时便于性能调优。
This study addresses the low utilization of modern GPU computing resources by systematically evaluating the performance, energy efficiency, and resource isolation characteristics of NVIDIA’s Multi-Process Service (MPS) and Multi-Instance GPU (MIG) technologies under concurrent application workloads. The experiments reveal a critical trade-off between MPS’s scheduling flexibility and MIG’s hardware-level isolation: MPS can improve performance by up to 30% and reduce energy consumption by approximately 20% in the absence of memory contention, yet suffers a 30% performance degradation under contention; MIG effectively mitigates resource contention but is constrained by its rigid configuration options and higher overhead. These findings provide empirical foundations for optimizing GPU co-execution strategies driven by application-specific workload characteristics.
This work systematically investigates the trade-offs between performance and programmability in CPU-GPU cooperative scheduling across discrete and unified memory architectures, with a focus on sparse conjugate gradient computations. Evaluations are conducted on both the NVIDIA GH200 Superchip—a platform featuring a unified memory architecture—and the discrete H100 PCIe system, comparing three memory management paradigms: explicit data copies, managed memory, and mapped memory. The study reveals that the GH200’s fused architecture substantially enhances the practicality of managed memory, enabling diverse hybrid task-partitioning strategies to achieve both high performance and programming simplicity. These findings underscore the significant impact of underlying memory architecture on the efficacy of cooperative scheduling approaches.
This work addresses the high energy consumption of traditional GPUs under the SIMT programming model, which stems from frequent register file accesses and complex control logic. The authors propose replacing the SIMD backend with a statically scheduled coarse-grained reconfigurable array (CGRA) that pipelines active threads and enables direct data transfer among processing elements, drastically reducing intermediate value accesses to registers. A novel p-graph program representation is introduced to decouple dynamic dependency edges, and in conjunction with double-buffered configuration memory, compile-time graph unrolling, and a temporal memory coalescing unit (TMCU), the design efficiently supports dynamic behavior while retaining static scheduling. Experimental results on the Rodinia benchmark suite show an average 68% reduction in register file accesses, 1.77–1.90× improvement in dynamic energy efficiency, 42.0%–45.9% lower power consumption, and performance comparable to that of NVIDIA Turing GPUs.
Current GPU programming models lack expressiveness for chiplet-level locality and synchronization, leading to redundant memory accesses and poor cache utilization when executing memory-intensive workloads such as large language model (LLM) inference on multi-chiplet GPUs. This work proposes Fleet, the first multi-level task programming model that explicitly exposes the chiplet hierarchy. Fleet introduces a chiplet-task abstraction that binds computation and data to specific chiplets and integrates persistent kernels, cooperative weight tiling, and per-chiplet scheduling to enable L2 cache reuse and efficient coordinated execution. Evaluated on an AMD MI350 running Qwen3-8B, Fleet reduces decoding latency by 1.3–1.5× for small batches and cuts HBM traffic by up to 37% under large batches, significantly improving L2 hit rates and achieving overall speedups of 1.27–1.30×.