🤖 AI Summary
Contemporary NVIDIA GPU microarchitectural research lags significantly, often relying on designs over fifteen years old. Method: This paper presents the first systematic reverse-engineering study of RTX-class GPU cores, uncovering their instruction scheduling policies, register file and cache hierarchies, memory pipeline characteristics, and hardware-software co-execution mechanisms. It proposes a stream-buffer-based instruction prefetcher and empirically demonstrates that software-managed dependency tracking outperforms traditional hardware scoreboard approaches. Contribution/Results: We develop a high-fidelity instruction-level simulator incorporating detailed register file caching and read-port modeling. Evaluated on the RTX A6000, it achieves a mean absolute percentage error (MAPE) of 13.98%, representing an 18.24% improvement over state-of-the-art simulators. Moreover, the model exhibits cross-generational generalizability, successfully transferring to the Turing architecture.
📝 Abstract
GPUs are the most popular platform for accelerating HPC workloads, such as artificial intelligence and science simulations. However, most microarchitectural research in academia relies on GPU core pipeline designs based on architectures that are more than 15 years old. This paper reverse engineers modern NVIDIA GPU cores, unveiling many key aspects of its design and explaining how GPUs leverage hardware-compiler techniques where the compiler guides hardware during execution. In particular, it reveals how the issue logic works including the policy of the issue scheduler, the structure of the register file and its associated cache, and multiple features of the memory pipeline. Moreover, it analyses how a simple instruction prefetcher based on a stream buffer fits well with modern NVIDIA GPUs and is likely to be used. Furthermore, we investigate the impact of the register file cache and the number of register file read ports on both simulation accuracy and performance. By modeling all these new discovered microarchitectural details, we achieve 18.24% lower mean absolute percentage error (MAPE) in execution cycles than previous state-of-the-art simulators, resulting in an average of 13.98% MAPE with respect to real hardware (NVIDIA RTX A6000). Also, we demonstrate that this new model stands for other NVIDIA architectures, such as Turing. Finally, we show that the software-based dependence management mechanism included in modern NVIDIA GPUs outperforms a hardware mechanism based on scoreboards in terms of performance and area.