Accelerating the Mitigation of LLM Inference Nondeterminism Across GPU Architectures

📅 2026-09-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
论文解决了LLM在不同GPU架构上推理结果不一致的问题,通过使用固定配置的融合升位GEMM内核方法,确保了跨架构的可重复性并提高了性能。
📝 Abstract
Large language model (LLM) outputs are expected to be reproducible under greedy decoding, yet in practice the same model, prompt, and software stack produce different outputs on different GPUs. The root cause is floating-point non-associativity combined with hardware-dependent kernel selection. Inference frameworks select different matrix-multiplication kernels on each architecture, with different parallel reduction orders and unspecified tensor-core arithmetic, and the resulting rounding differences can flip output tokens. Existing solutions have imperfect cross-architecture reproducibility and incur a significant performance penalty. We present a solution employing a set of fixed-configuration fused-upcast GEMM kernels that load 16-bit weights from memory, upcast them to FP32 in registers, and accumulate with IEEE-754 arithmetic in a reduction order that is a pure function of the problem shape and is therefore independent of the device, its SM count, or kernel scheduling. By fixing the floating-point reduction order as a function of problem shape alone, every GPU runs the same operation sequence, so cross-architecture reproducibility of the linear layers reduces to correct IEEE-754 arithmetic rather than to rounding differences staying below a tie-flip threshold. We confirm our solution's linear-layer outputs are bitwise identical across NVIDIA Ampere, Ada, and Hopper GPUs, while running $1.17$ to $3.1\times$ faster end-to-end than the state-of-the-art solution and cutting weight-memory traffic in half.
Problem

Research questions and friction points this paper is trying to address.

LLM
inference nondeterminism
GPU architectures
floating-point non-associativity
kernel selection
Innovation

Methods, ideas, or system contributions that make the work stand out.

fixed-configuration fused-upcast GEMM kernels
cross-architecture reproducibility
floating-point reduction order
IEEE-754 arithmetic
🔎 Similar Papers
No similar papers found.