Delta-Matching: Closing the Final Gap of Native 8-bit Training for LLMs

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the gradient distortion and training instability in native FP8 large model training caused by forward-backward inconsistency within the attention mechanism. We propose an end-to-end, architecture-agnostic native block-scaled FP8 training scheme that requires no architectural modifications. Theoretically, we demonstrate that Delta-Matching restores the zero-row-sum invariant of Softmax gradients, which, combined with strategies such as QK normalization, effectively eliminates quantization errors. Extensive experiments show that our approach achieves convergence loss and downstream performance consistent with BF16/FP32 mixed-precision training across diverse architectures and model scales. Furthermore, we have open-sourced the complete implementation code and model weights to facilitate broader adoption.
📝 Abstract
Reliable FP8 attention remains a barrier to fully native 8-bit large language model training. We derive how forward-backward inconsistencies produce stale delta and empirically show how it distorts training dynamics. Our stale-delta hybrid runs show a modest loss gap at 569M parameters but substantial loss increases and downstream degradation at 1.67B and 5.29B. QK normalization, NoPE (no positional encoding), and lower-learning-rate context extension mitigate or delay degradation without eliminating it. This pattern suggests accumulated optimization error that smaller models and short runs can conceal. We propose Delta-Matching, proving that it restores the softmax gradient's zero-row-sum invariant under the stated numerical assumptions. It enables native block-scaled FP8 in every forward and backward attention-core matmul without architectural changes, smaller global batches, or auxiliary forward outputs. Across tested architectures, scales, and training stages, Delta-Matching matches BF16/FP32 mixed-precision training loss and overall downstream performance. We will release our implementation, trained models, and data recipes.
Problem

Research questions and friction points this paper is trying to address.

FP8 training
large language models
attention mechanism
stale delta
training degradation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Native 8-bit Training
Delta-Matching
FP8 Attention
Stale Delta
Zero-Row-Sum Invariant