Pretraining Transformers with Quantized Softmax in Attention

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that quantizing Softmax during low-precision Transformer pre-training disrupts both forward computation and backward gradient propagation, while existing calibration strategies struggle to balance efficiency and accuracy. To overcome this, we propose a quantization scheme based on K-interval attention approximation, systematically optimizing grid calibration, rounding methods, and straight-through estimator (STE) placement. Specifically, we introduce a fixed-window calibration combined with a post-normalization STE, accompanied by rigorously derived backpropagation rules. Evaluated on a 124-million-parameter model, our approach incurs only a 0.004-nat increase in validation loss at K=16, achieving near-full-precision training performance with minimal overhead. This work provides a reliable paradigm for efficient quantized pre-training.
📝 Abstract
Low-precision Transformer systems increasingly quantize attention matrix multiplications, while softmax often remains at higher precision. During pretraining, an approximate softmax changes the gradients that train the model as well as its forward computation. We study this interaction with K-interval attention, which approximates the exponential using K+1 grid values. We vary per-row grid calibration, interpolation versus hard rounding, and the placement of a straight-through surrogate relative to normalization. We derive the corresponding backward rules, including calibration derivatives, and compare these choices in pretraining experiments matched on model, data, and optimizer. Detaching the row extrema leaves the forward computation unchanged but produces a delayed increase in validation loss. With hard rounding at K=4, min-max calibration and a pre-normalization surrogate incur a large loss gap; changing either choice substantially reduces it. At 124M parameters and 2.5B training tokens, fixed-window calibration with a post-normalization surrogate yields a validation loss gap of +0.019 nats relative to softmax at K=4, and with a pre-normalization surrogate yields +0.004 nats at K=16.
Problem

Research questions and friction points this paper is trying to address.

quantized softmax
Transformer pretraining
low-precision attention
approximate softmax
gradient interaction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Quantized Softmax
K-interval Attention
Straight-Through Estimator
Low-Precision Transformer
Pretraining
S
Shangzhen Zhu
University of Illinois Urbana-Champaign
M
Muyan Hu
University of Illinois Urbana-Champaign
Tomasz Kozlowski
Tomasz Kozlowski
University of Illinois Urbana-Champaign