🤖 AI Summary
This study addresses the computational bottleneck of Softmax exponentiation in large language model inference by proposing Rowmax-PoT, a method that approximates exponential operations using powers of two via coarse-grained logarithmic weight representations anchored to row maxima. Motivated by the finding that resolution budget allocation matters more than absolute magnitude, we implement hardware-specialized kernels within the FlashAttention-4 framework, leveraging NVIDIA B200 architecture and FP8/BF16 mixed-precision Tensor Cores. Experiments demonstrate that on the B200 platform, forward-pass throughput for 8K sequences improves by 12.4% and 25.8% for causal and non-causal attention, respectively, while energy consumption for 16K sequences decreases by 8.4%. Notably, perplexity incurs only a marginal increase of 0.09%–0.49%, achieving substantial energy efficiency gains with minimal precision cost.
📝 Abstract
On NVIDIA Blackwell B200, tensor-core throughput outpaces special-function exponential throughput by more than two orders of magnitude, exposing exponential evaluation in fused attention kernels. A pretrained Transformer, however, may not need it evaluated accurately at every element. We characterize what a pretrained model does need by approximating softmax at inference in ten frozen decoder-only models (0.5B-72B). The number of positions the softmax map assigns probability to and within-row resolution can be cut substantially, yet uniform weighting of the same positions is damaging. Where a fixed resolution budget is placed matters as much as its size, with resolution near the row maximum consistently favored. Perturbations matched on scalar distortion produce model-dependent responses of opposite sign. These findings motivate Rowmax-PoT, a coarse logarithmic weight representation anchored at each row maximum, and Rowmax-H15, its hardware specialization in FlashAttention-4. On B200, the patched FP8 attention forward is 12.4% faster at causal 8K and 25.8% faster at non-causal 8K in host-side call-latency measurements; board energy per forward falls by 8.4% at causal 16K. Measured separately on the BF16 kernel path at 2K, Rowmax-H15 increases perplexity by 0.091-0.492% across five models from three families.