Stochastic Rounding in Low-Precision Transformer Inference: A Variable-Precision Emulation Study of a Small GPT-2
This study addresses the unclear applicability of stochastic versus nearest rounding in low-precision Transformer inference. We propose a variable-precision simulation method and probabilistic forward error bound analysis, revealing for the first time the differential sensitivity of network layers to rounding errors. Building on second-order loss decomposition theory, we design a layer-wise mixed rounding strategy and a variable-precision stochastic rounding algorithm. Evaluated on DistilGPT-2, our approach achieves a mixed-precision configuration that reduces perplexity to 1.10× that of full precision, yielding a 28% improvement over uniform nearest rounding. This work provides an efficient quantization scheme for low-precision inference.