Analyzing Modern NVIDIA GPU cores

📅 2025-03-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Contemporary NVIDIA GPU microarchitectural research lags significantly, often relying on designs over fifteen years old. Method: This paper presents the first systematic reverse-engineering study of RTX-class GPU cores, uncovering their instruction scheduling policies, register file and cache hierarchies, memory pipeline characteristics, and hardware-software co-execution mechanisms. It proposes a stream-buffer-based instruction prefetcher and empirically demonstrates that software-managed dependency tracking outperforms traditional hardware scoreboard approaches. Contribution/Results: We develop a high-fidelity instruction-level simulator incorporating detailed register file caching and read-port modeling. Evaluated on the RTX A6000, it achieves a mean absolute percentage error (MAPE) of 13.98%, representing an 18.24% improvement over state-of-the-art simulators. Moreover, the model exhibits cross-generational generalizability, successfully transferring to the Turing architecture.

Technology Category

Machine Learning: Hardware-aware MLCognitive Modeling & Cognitive Systems: Simulating Human BehaviorSearch and Optimization: Sampling/Simulation-based Search

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsSystems and Infrastructure for Web, Mobile and WoT: Web performance, measurement, and characterizationWeb Mining and Content Analysis: Web data generation and simulation
📝 Abstract
GPUs are the most popular platform for accelerating HPC workloads, such as artificial intelligence and science simulations. However, most microarchitectural research in academia relies on GPU core pipeline designs based on architectures that are more than 15 years old. This paper reverse engineers modern NVIDIA GPU cores, unveiling many key aspects of its design and explaining how GPUs leverage hardware-compiler techniques where the compiler guides hardware during execution. In particular, it reveals how the issue logic works including the policy of the issue scheduler, the structure of the register file and its associated cache, and multiple features of the memory pipeline. Moreover, it analyses how a simple instruction prefetcher based on a stream buffer fits well with modern NVIDIA GPUs and is likely to be used. Furthermore, we investigate the impact of the register file cache and the number of register file read ports on both simulation accuracy and performance. By modeling all these new discovered microarchitectural details, we achieve 18.24% lower mean absolute percentage error (MAPE) in execution cycles than previous state-of-the-art simulators, resulting in an average of 13.98% MAPE with respect to real hardware (NVIDIA RTX A6000). Also, we demonstrate that this new model stands for other NVIDIA architectures, such as Turing. Finally, we show that the software-based dependence management mechanism included in modern NVIDIA GPUs outperforms a hardware mechanism based on scoreboards in terms of performance and area.
Problem

Research questions and friction points this paper is trying to address.

Reverse engineer modern NVIDIA GPU cores' design
Analyze instruction prefetcher and register file impact
Improve simulation accuracy for HPC workloads
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reverse engineers modern NVIDIA GPU cores
Analyzes instruction prefetcher with stream buffer
Models microarchitectural details for accuracy
🔎 Similar Papers
No similar papers found.
R
Rodrigo Huerta
Universitat Politècnica de Catalunya, Barcelona, Spain
Mojtaba Abaie Shoushtary
Mojtaba Abaie Shoushtary
Polytechnic University of Catalonia
Computer ArchitectureGPGPUArtificial IntelligenceMachine LearningDeep Learning
J
José-Lorenzo Cruz
Universitat Politècnica de Catalunya, Barcelona, Spain
A
A. González
Universitat Politècnica de Catalunya, Barcelona, Spain