Stochastic Rounding in Low-Precision Transformer Inference: A Variable-Precision Emulation Study of a Small GPT-2

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear applicability of stochastic versus nearest rounding in low-precision Transformer inference. We propose a variable-precision simulation method and probabilistic forward error bound analysis, revealing for the first time the differential sensitivity of network layers to rounding errors. Building on second-order loss decomposition theory, we design a layer-wise mixed rounding strategy and a variable-precision stochastic rounding algorithm. Evaluated on DistilGPT-2, our approach achieves a mixed-precision configuration that reduces perplexity to 1.10× that of full precision, yielding a 28% improvement over uniform nearest rounding. This work provides an efficient quantization scheme for low-precision inference.
📝 Abstract
Should low-precision transformer inference use stochastic rounding (SR) or round-to-nearest (RN)? The answer depends on where in the network you look. We isolate this effect by holding the numerical format fixed and varying only the rounding rule at individual operation sites. To enable experiments at freely chosen precisions, we extend the PRISM vectorized rounding library to arbitrary virtual precision via a variable-precision stochastic rounding (VPSR) algorithm, proving that the rounding decision is evaluated exactly in hardware floating point. We develop two analyses providing complementary insight into this site-level trade-off. First, a probabilistic forward-error bound for linear projections shows that SR's error envelope grows as $O(\sqrt{n} u)$ in reduction length $n$, versus $O(n u)$ for RN, a gap that widens rapidly at low precision and is most pronounced in the long multilayer perceptron (MLP) down-projection. Second, a second-order decomposition of expected cross-entropy loss change at the output softmax into signed drift, drift curvature, and a Fisher-weighted variance penalty reveals why the two sites behave oppositely: MLP noise is predominantly a uniform logit shift to which softmax is invariant, so SR's variance is largely discounted; head noise is non-uniform across the vocabulary and is not. On DistilGPT-2 at $t=6$ significand bits, observations match theory: SR in the MLP raises perplexity to 1.15x the full-precision reference, versus 2.21x for RN. At the language-model head, the ordering reverses because SR introduces non-uniform variance, whereas deterministic RN carries none. In a mixed-precision configuration (MLP output at $t=6$), assigning SR to the MLP and RN to the head brings perplexity within 1.10x of the full-precision reference, a 28% reduction over matched-bit RN.
Problem

Research questions and friction points this paper is trying to address.

Stochastic Rounding
Low-Precision Inference
Transformer
Round-to-Nearest
Variable Precision
Innovation

Methods, ideas, or system contributions that make the work stand out.

Stochastic Rounding
Variable-Precision Emulation
Low-Precision Inference
Transformer
Mixed-Precision
🔎 Similar Papers
No similar papers found.