The Missing Fourth Term for the Emulation Tensor Memory Equilibrium (TME) Model: The Residue Deconstruction Cost

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the bias in performance evaluation of memory-bound operators caused by existing FP8 simulation models, which neglect the overhead of input residual decomposition. Based on the Tensor-Memory Equilibrium (TME) model, this work incorporates a fourth cost term to quantify, for the first time, the residual decomposition cost on SIMT pipelines and establish a closed-form theoretical lower bound. By integrating the Ozaki Scheme II algorithm with cuBLAS path calibration, operational intensity thresholds are derived to precisely delineate the limits of simulated acceleration. Analysis targeting NVIDIA B300 hardware reveals a threshold of 0.56 FLOP/B, demonstrating that operators such as GEMV achieve only 0.3–0.9× native performance. These findings effectively correct prior overestimations of performance in low-batch scenarios.
📝 Abstract
The Tensor-Memory Equilibrium (TME) model of "FP8 is All You Need (Part 1)" calculates the execution time of Ozaki Scheme II emulation of fp64 as the maximum of a tensor-core term and a High-Bandwidth Memory (HBM) traffic term, plus a per-output reconstruction term. However, it omits the per-input deconstruction cost: every streamed fp64 operand must be scaled, rounded, and reduced modulo each of the $r$ moduli on SIMT pipes before any matrix multiply can issue. In this note we add this fourth term, calibrate its constant from the cuBLAS emulation path, and derive a closed-form operational-intensity threshold $\mathrm{OI}^{*} = c_q r P_{\mathrm{fp64}}/(8P_{\mathrm{int}})$ below which emulation cannot match native fp64 regardless of tensor-core throughput. On the NVIDIA B300 GPU the threshold is $\mathrm{OI}^{*}\approx 0.56$ FLOP/B. As a result, GEMV, SpMV, and low-batch GEMV, which are the memory-bound kernels the original paper claims to accelerate, are limited to 0.3-0.9x of native performance, and the 7-point stencil to 1.8x rather than the claimed 3.1x. Dense GEMM is unaffected as expected. We also show that precomputing and storing the residues moves the same cost into the bandwidth term, and we state the instruction count that an implementation would have to achieve to invalidate the bound.
Problem

Research questions and friction points this paper is trying to address.

Tensor-Memory Equilibrium
fp64 emulation
deconstruction cost
operational intensity
performance modeling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Tensor-Memory Equilibrium
Residue Deconstruction Cost
Operational Intensity Threshold
FP8 Emulation
Performance Modeling