Greedy Decoding Is Not Precision-Invariant: Cross-Precision Output Divergence in LLM Inference

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究解决了大语言模型贪婪解码在不同精度下输出不一致的问题,通过选择性FP32 LM头部重计算方法提高了输出一致性。
📝 Abstract
Greedy decoding from large language models is commonly treated as deterministic. We show it is not precision-invariant: the same model, prompt, and decoding algorithm produce different outputs in BF16 versus FP16 on identical hardware. Across our evaluations of six models (1.1B-7B parameters, four families; divergence additionally characterised at 12B) and three benchmarks, 49-100\% of prompts diverge; a single token flip often cascades into trajectory-level divergence. We develop an empirical error-propagation analysis and find that 22 layers of accumulated body error do not distinguish flipping from non-flipping steps; the outcome depends primarily on the top-two logit margin at the LM head relative to the directional perturbation between the top-two candidates. The analysis makes five testable predictions about intervention outcomes, including that applying more FP32 compute (broader scope) makes agreement worse. The experiments match all five predictions. The best-performing low-overhead intervention we evaluate, selective FP32 LM head recomputation, triggered only when the margin falls below a threshold, delivers +22-36 pp exact agreement on A10G (+12-21 pp on L4 and A100) at less than 4\% latency overhead in low-batch (batch size <=4) single-stream inference. We map the applicability boundary across six models and four batch sizes, and hypothesise that training-time precision stability is a determining factor. The method is a partial mitigation rather than a universal determinism guarantee: its benefit vanishes when body-originated error dominates, including at batch size >=8 and under end-to-end FP8 in our tests.
Problem

Research questions and friction points this paper is trying to address.

greedy decoding
precision-invariance
output divergence
large language models
BF16 vs FP16
Innovation

Methods, ideas, or system contributions that make the work stand out.

Greedy Decoding
Precision Invariance
Error Propagation Analysis
Selective FP32 Recomputation
Low-Overhead Intervention
🔎 Similar Papers
No similar papers found.