Approximating Softmax in Pretrained LLMs: Model Sensitivity and Kernel Acceleration

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the computational bottleneck of Softmax exponentiation in large language model inference by proposing Rowmax-PoT, a method that approximates exponential operations using powers of two via coarse-grained logarithmic weight representations anchored to row maxima. Motivated by the finding that resolution budget allocation matters more than absolute magnitude, we implement hardware-specialized kernels within the FlashAttention-4 framework, leveraging NVIDIA B200 architecture and FP8/BF16 mixed-precision Tensor Cores. Experiments demonstrate that on the B200 platform, forward-pass throughput for 8K sequences improves by 12.4% and 25.8% for causal and non-causal attention, respectively, while energy consumption for 16K sequences decreases by 8.4%. Notably, perplexity incurs only a marginal increase of 0.09%–0.49%, achieving substantial energy efficiency gains with minimal precision cost.
📝 Abstract
On NVIDIA Blackwell B200, tensor-core throughput outpaces special-function exponential throughput by more than two orders of magnitude, exposing exponential evaluation in fused attention kernels. A pretrained Transformer, however, may not need it evaluated accurately at every element. We characterize what a pretrained model does need by approximating softmax at inference in ten frozen decoder-only models (0.5B-72B). The number of positions the softmax map assigns probability to and within-row resolution can be cut substantially, yet uniform weighting of the same positions is damaging. Where a fixed resolution budget is placed matters as much as its size, with resolution near the row maximum consistently favored. Perturbations matched on scalar distortion produce model-dependent responses of opposite sign. These findings motivate Rowmax-PoT, a coarse logarithmic weight representation anchored at each row maximum, and Rowmax-H15, its hardware specialization in FlashAttention-4. On B200, the patched FP8 attention forward is 12.4% faster at causal 8K and 25.8% faster at non-causal 8K in host-side call-latency measurements; board energy per forward falls by 8.4% at causal 16K. Measured separately on the BF16 kernel path at 2K, Rowmax-H15 increases perplexity by 0.091-0.492% across five models from three families.
Problem

Research questions and friction points this paper is trying to address.

Softmax approximation
Pretrained LLMs
Hardware acceleration
Attention kernel
Model sensitivity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Softmax approximation
Rowmax-PoT
FlashAttention-4
NVIDIA Blackwell B200
Kernel acceleration
🔎 Similar Papers
No similar papers found.