π€ AI Summary
This work addresses the significant degradation of FP64 performance in next-generation AI-optimized GPUs (e.g., NVIDIA B300), which undermines their suitability for scientific computing. The authors propose a novel paradigm that reconstructs all Berkeley βdwarfsβ kernel functions exclusively using FP8 tensor-core matrix multiplication as the sole computational primitive, augmented by Ozaki Scheme II based on the Chinese Remainder Theorem, register-level fusion, compensated reduction, and FFT-based analysis. Crucially, double-precision accuracy is preserved through finite-width integer accumulation. The study establishes a five-layer theoretical framework, effecting the first paradigm shift from hardware-dependent FP64 to compositionally derived FP64 via FP8 operations, and introduces Tensor-Memory Equilibrium (TME), an extension of the Roofline model, for performance evaluation. Experiments demonstrate that the approach recovers near-H100 FP64 performance across diverse kernels while retaining memory-bound characteristics, thereby validating its theoretical feasibility.
π Abstract
Conventional HPC dogma holds that native hardware FP64 silicon is the irreducible foundation of scientific computing -- the"holy grail"of double-precision simulation. This paper argues the dogma is wrong: on AI-optimised GPUs of the B300 generation and beyond, abundant FP8 tensor throughput combined with the Chinese Remainder Theorem-based Ozaki Scheme II recovers memory-roof execution at full FP64 accuracy across the canonical HPC kernel spectrum. NVIDIA's Blackwell Ultra (B300) collapses native FP64 to ~1.3 TFLOPS -- a 31x regression from the B200 -- rendering even memory-bound kernels (SpMV, GEMV, stencils) compute-bound. We make four contributions. First, a unified analytic model, the Tensor-Memory Equilibrium (TME) model, augmenting the Roofline with a compute multiplier alpha, a bandwidth multiplier beta, and a reconstruction latency gamma. Second, we identify register-level fusion as the mechanism driving beta ->1, making emulation essentially free behind the memory wall. Third, we project that Ozaki II vaults emulated FP64 from the ~1 TFLOPS native floor to ~500 TFLOPS (B300) and ~400 TFLOPS (Rubin R200), exceeding even B200's native FP64 ceiling by over an order of magnitude in the compute-bound regime while matching the memory roof in the bandwidth-bound regime. Fourth, against an H100 baseline, Ozaki II matches or exceeds H100 on every workload studied, versus the up-to-50x regression that B300 native FP64 imposes. Combined with a companion FFT analysis (Kulisch fixed-point reconstruction on the surviving INT32 pipe) and FP32+Kahan reductions reported in the companion Part(2) paper, every surveyed kernel class on B300 reaches the memory roof at full FP64. The evidence supports the title's claim: FP8, with Ozaki II and Kulisch escape routes, is all one needs for production HPC; native FP64 silicon is no longer the holy grail it has been taken to be.