FP8 is All You Need (Part 1): Debunking Hardware FP64 as the HPC Holy Grail

πŸ“… 2026-05-28
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the significant degradation of FP64 performance in next-generation AI-optimized GPUs (e.g., NVIDIA B300), which undermines their suitability for scientific computing. The authors propose a novel paradigm that reconstructs all Berkeley β€œdwarfs” kernel functions exclusively using FP8 tensor-core matrix multiplication as the sole computational primitive, augmented by Ozaki Scheme II based on the Chinese Remainder Theorem, register-level fusion, compensated reduction, and FFT-based analysis. Crucially, double-precision accuracy is preserved through finite-width integer accumulation. The study establishes a five-layer theoretical framework, effecting the first paradigm shift from hardware-dependent FP64 to compositionally derived FP64 via FP8 operations, and introduces Tensor-Memory Equilibrium (TME), an extension of the Roofline model, for performance evaluation. Experiments demonstrate that the approach recovers near-H100 FP64 performance across diverse kernels while retaining memory-bound characteristics, thereby validating its theoretical feasibility.
πŸ“ Abstract
Conventional HPC dogma holds that native hardware FP64 silicon is the irreducible foundation of scientific computing -- the"holy grail"of double-precision simulation. This paper argues the dogma is wrong: on AI-optimised GPUs of the B300 generation and beyond, abundant FP8 tensor throughput combined with the Chinese Remainder Theorem-based Ozaki Scheme II recovers memory-roof execution at full FP64 accuracy across the canonical HPC kernel spectrum. NVIDIA's Blackwell Ultra (B300) collapses native FP64 to ~1.3 TFLOPS -- a 31x regression from the B200 -- rendering even memory-bound kernels (SpMV, GEMV, stencils) compute-bound. We make four contributions. First, a unified analytic model, the Tensor-Memory Equilibrium (TME) model, augmenting the Roofline with a compute multiplier alpha, a bandwidth multiplier beta, and a reconstruction latency gamma. Second, we identify register-level fusion as the mechanism driving beta ->1, making emulation essentially free behind the memory wall. Third, we project that Ozaki II vaults emulated FP64 from the ~1 TFLOPS native floor to ~500 TFLOPS (B300) and ~400 TFLOPS (Rubin R200), exceeding even B200's native FP64 ceiling by over an order of magnitude in the compute-bound regime while matching the memory roof in the bandwidth-bound regime. Fourth, against an H100 baseline, Ozaki II matches or exceeds H100 on every workload studied, versus the up-to-50x regression that B300 native FP64 imposes. Combined with a companion FFT analysis (Kulisch fixed-point reconstruction on the surviving INT32 pipe) and FP32+Kahan reductions reported in the companion Part(2) paper, every surveyed kernel class on B300 reaches the memory roof at full FP64. The evidence supports the title's claim: FP8, with Ozaki II and Kulisch escape routes, is all one needs for production HPC; native FP64 silicon is no longer the holy grail it has been taken to be.
Problem

Research questions and friction points this paper is trying to address.

FP8
FP64
scientific computing
tensor cores
high-performance computing
Innovation

Methods, ideas, or system contributions that make the work stand out.

FP8 tensor cores
Ozaki Scheme II
Tensor-Memory Equilibrium
double-precision emulation
HPC dwarfs
πŸ”Ž Similar Papers
No similar papers found.