Measuring and Reducing Cross-Vendor Mismatch in Language Models

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses numerical discrepancies of identical language models across GPU vendors, attributing the root cause to divergent matrix accumulation orders. To mitigate this, the authors employ mixed-precision strategies (FP32/FP16/BF16), LoRA adaptation, and knowledge distillation for both dense and Mixture-of-Experts (MoE) architectures, alongside multi-granularity evaluation metrics including bit-level equivalence and logit divergence. Key findings reveal that retaining BF16 precision exclusively in MLP layers yields primary gains, while single-precision upcasting disrupts MoE expert routing, exposing the limitations of single-metric evaluations. Empirically, partial high-precision computation reduces logit errors by 94% with only a 30% latency overhead for dense models, whereas cross-vendor distillation induces significant deviations from MMLU benchmarks.
📝 Abstract
Running the same language model on different graphics processing unit (GPU) vendors can produce different logits, even when the model weights and inputs are the same. We analyze cross-vendor mismatch in two dense and two mixture-of-experts (MoE) models with five metric families, namely bitwise equality, logit differences, top-K consistency, token agreement, and task accuracy. We trace one source of the mismatch to accumulation order inside vendors'matrix instructions. Upcasting to FP32 reduces the dense model's logit error by 43% at three times the runtime, yet keeping only the MLPs in BF16 retains 94% of this gain at 1.3 times the runtime, so most of the cost of full upcasting buys little. In the MoE models, FP32 and FP16 both lower the probability error but raise the logit error and change expert selection, and FP16 fails in the dense model. An output-head low-rank adapter (LoRA) does not help either, since the final hidden state does not predict the mismatch. The mismatch also carries into training. With every seed fixed, a student distilled from a teacher running on AMD answers 431 MMLU questions differently from one distilled from the same teacher on NVIDIA. Under FP32 upcasting, bitwise equality barely changes while the output distributions move most of the way to the reference, so judging cross-vendor agreement by a single measure misreads both its cost and its gains. Code is available at https://github.com/crova-project/crova.
Problem

Research questions and friction points this paper is trying to address.

cross-vendor mismatch
language models
GPU
logit differences
knowledge distillation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-Vendor Mismatch
Mixed Precision Upcasting
Mixture-of-Experts (MoE)
Knowledge Distillation
Accumulation Order
🔎 Similar Papers