The Silent Freeze: Predicting When Low-Precision Training Stops Learning

๐Ÿ“… 2026-07-09
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the issue of silent, irreversible training stagnation in low-precision optimization, where weight updates are truncated when rounding errors fall below half a unit in the last place (ULP). The study presents and validates, for the first time, a deterministic condition for weight freezing in low-precision training: by leveraging the high-precision (fp32) training trajectory and the mantissa length of the target floating-point format (e.g., fp8 or bf16), the exact iteration at which each weight coordinate freezes can be predicted a prioriโ€”without executing actual low-precision training. Using a mantissa-truncation simulator combined with stochastic rounding, the method accurately forecasts freezing points on models such as GPT and GPT-2 with an error of at most four steps, revealing that weight freezing is the primary cause of validation loss plateaus. Stochastic rounding effectively mitigates this phenomenon, and the approach generalizes across diverse models, tasks, and floating-point formats.
๐Ÿ“ Abstract
Training in reduced floating-point precision can silently halt learning: when a gradient-descent weight update falls below half the unit in the last place (ULP) of the weight, it rounds away and that coordinate freezes while its gradient is still nonzero. The freeze is deterministic, governed by a per-coordinate half-ULP condition, and predictable from a high-precision trajectory and the target mantissa length alone, without low-precision data. In a small GPT trained under the standard AdamW-plus-cosine recipe with bf16-equivalent stored weights, training proceeds normally and then permanently freezes just past mid-run, within four steps of the a-priori prediction. In a $124$-million-parameter GPT-2 transformer whose weights are constrained to the $8$-bit floating-point grid after every optimizer step, with no master weights, the dense weights freeze at initialization in both fp8 formats -- predicted \emph{a priori} from an fp32 reference -- and validation loss plateaus while full precision keeps improving. Stochastic rounding removes the persistent freeze, and the same reference predicts that too. The condition transfers across frozen-feature regression, a mantissa-truncation emulator spanning $128\times$ in precision, small networks, and a CNN on MNIST: a computable axis of low-precision training, not diffuse noise.
Problem

Research questions and friction points this paper is trying to address.

low-precision training
silent freeze
gradient descent
floating-point precision
weight update
Innovation

Methods, ideas, or system contributions that make the work stand out.

silent freeze
low-precision training
half-ULP condition
predictable freezing
stochastic rounding
๐Ÿ”Ž Similar Papers
No similar papers found.