mixed-precision training

Designs, implements, and evaluates numerical precision schemes for training and inference of machine learning models in which weights, activations, optimizer moments, and compute kernels use different floating- or integer-bit widths (e.g., bf16 weights with fp32 moments) to reduce runtime and energy while maintaining model accuracy. This includes sensitivity-guided and inter-/intra-layer bit-allocation algorithms, mixed-precision programming and low-precision kernels, precision-aware optimizers and validation workflows (benchmarking throughput vs accuracy, aggregating metrics such as perplexity and reasoning signals), and techniques to prevent numerical cascades and preserve stability.

mixed-precisiontraining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-2.8
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$218K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Standalone 16-bit Neural Network Training: Missing Study for Hardware-Limited Deep Learning Practitioners

May 18, 2023
JY
Juyoung Yun
🏛️ Stony Brook University | State University of New York

Whether 16-bit floating-point (FP16) arithmetic alone can robustly support end-to-end neural network training under hardware resource constraints remains a long-standing, systematically unverified question. Method: This paper introduces the first unified framework for classification-tolerant error analysis, integrating theoretical floating-point error modeling with large-scale empirical validation, and implements fully FP16 forward and backward propagation—without any FP32 “master weights” or loss scaling. Contribution/Results: Rigorous evaluation on CIFAR-10, CIFAR-100, and ImageNet demonstrates that pure FP16 training achieves accuracy parity with FP32 or mixed-precision baselines (accuracy deviation ≤ ±0.1%), while accelerating training by 1.8×. All experiments deploy seamlessly on mainstream GPUs with zero code modification. This work bridges a critical gap between theoretical guarantees and practical feasibility of low-precision training.

Enable efficient neural network training on limited hardwareExplore conditions for 16-bit precision approximationValidate 16-bit precision matches 32-bit accuracy

Algorithms and data structures for automatic precision estimation of neural networks

Sep 29, 2025
IV
Igor V. Netay
🏛️ Joint Stock "Research and production company "Kryptonite" | Institute for Information Transmission Problems, Russian Academy of Sciences

Neural networks suffer from accumulated rounding errors in floating-point arithmetic, causing deviations between actual behavior and mathematical expectations—thereby compromising reliability and interpretability of inference and training. This paper introduces the first automated precision estimation method tailored for deep learning frameworks: it employs lightweight, differentiable data structures and algorithms to enable real-time error propagation tracking during both training and inference, balancing high-fidelity numerical modeling with computational efficiency while seamlessly integrating into mainstream neural network libraries. Its core contribution lies in systematizing and automating floating-point error analysis, enabling end-to-end numerical error monitoring. Extensive experiments across diverse models and tasks demonstrate the method’s broad applicability; they reveal pervasive and significant numerical distortions in most neural networks, underscoring the critical role of precision awareness in ensuring model robustness and trustworthiness.

Addressing computational inaccuracies affecting inference and trainingAutomatically estimating floating-point precision in neural networksEnsuring reliability and interpretability of neural network results

Bitwidth-Specific Logarithmic Arithmetic for Future Hardware-Accelerated Training

Oct 19, 2025
HH
Hassan Hamad
🏛️ University of Southern California

To address the high computational cost and low hardware efficiency of floating-point operations in deep learning training, this paper proposes a hardware-aware low-precision logarithmic fixed-point training method tailored for accelerators. The approach innovatively incorporates a bit-width–aware mechanism into logarithmic addition approximation, jointly optimizing piecewise linear approximation and simulated annealing to achieve Pareto-optimal trade-offs between accuracy and hardware overhead. Leveraging the logarithmic number system (LNS) and bit-accurate C++ simulation, end-to-end training is realized using 12-bit integer arithmetic. Experimental results on VGG-11 and VGG-16 demonstrate accuracy comparable to 32-bit floating-point training, while reducing multiply-accumulate (MAC) unit area by 32.5% and energy consumption by 53.5%. This work establishes a novel hardware–software co-design paradigm for low-precision deep learning training.

Enhancing low-precision logarithmic fixed-point training for hardware acceleratorsOptimizing bitwidth-specific approximations for logarithmic arithmetic operationsReducing area and energy in training with minimal accuracy loss

On Stochastic Rounding with Few Random Bits

Apr 29, 2025
AF
Andrew Fitzgibbon
🏛️ Graphcore

This work addresses the reliance of stochastic rounding (SR) on high-entropy random bits in low-precision (e.g., FP16) and mixed-precision computing, revealing a previously overlooked systematic bias induced by finite-bit random sources (FBSRs)—a bias invisible under infinite-precision theory yet detrimental to numerical reliability. Through rigorous error modeling, floating-point rounding analysis, and empirical training (e.g., ResNet-18), we quantitatively characterize the bias magnitude across multiple FBSR schemes for the first time, demonstrating up to a 1.2% degradation in training accuracy. Our study extends the reliability assessment framework for low-precision computation by explicitly incorporating FBSR-induced bias as a critical dimension. We further propose a low-bit SR implementation framework that explicitly controls bias while preserving efficiency, and release open-source, reproducible code. This work bridges theoretical SR analysis and practical low-precision system design, enabling more robust and predictable stochastic quantization in deep learning accelerators.

Examining bias in few-bit stochastic rounding implementationsImpact of rounding biases in low-precision machine learningReducing random bits while maintaining stochastic rounding benefits

Conventional FP32 multipliers incur substantial hardware overhead and poor energy efficiency, hindering high-performance CNN inference. Method: This paper proposes a co-optimization framework for deploying error-bounded approximate FP32 multipliers *within* convolution kernels, targeting CNN inference. Leveraging the IEEE 754 standard, it employs approximate compressors for significand multiplication and integrates NSGA-II—a multi-objective genetic algorithm—to jointly optimize multiplier type selection, placement, and composition order, thereby balancing accuracy and hardware efficiency. Contribution/Results: Evaluated across multiple CNN models, the approach maintains >99% of the original accuracy while reducing multiplier area and power consumption by 32.7% and 28.4% on average, respectively. This yields significant improvements in inference energy efficiency and throughput, providing a deployable, architecture-level solution for high-precision approximate computing.

Balancing accuracy and hardware efficiency using genetic algorithmsDesigning approximate FP32 multipliers to reduce hardware costsOptimizing CNN performance with interleaved approximate multipliers

Latest Papers

What's happening recently
View more

This work addresses the prevalent overuse of double-precision floating-point numbers in numerical programs by proposing an automated mixed-precision tuning methodology. The approach supports user-defined, non-standard low-precision floating-point formats with customizable exponent and mantissa bit-widths, and integrates numerical validation with systematic search within a unified framework to automatically generate program variants that meet prescribed accuracy constraints. Leveraging the PROMISE tool and containerized parallel benchmarking, the method demonstrates that numerous variables across a range of numerical applications and the Rodinia benchmark suite can be safely downgraded in precision. This reduction yields significant improvements in performance while simultaneously decreasing memory consumption and energy usage, all without compromising numerical accuracy.

floating-point formatsmixed-precisionnumerical accuracy

Existing large language models struggle to balance performance and accuracy under uniform low-precision floating-point quantization. This work proposes dMX, the first differentiable mixed-precision quantization framework for the OCP-defined MXFP format. By compressing discrete per-layer bit-width search into a single continuous learnable offset and integrating temperature-annealing scheduling with target-aware regularization, dMX enables a smooth transition from training to hardware deployment. Experiments on Llama, Qwen3, and SmolLM2 demonstrate that dMX consistently outperforms KL-divergence-based heuristic methods across WikiText-2 perplexity and four zero-shot reasoning tasks, achieving Pareto-optimal trade-offs between accuracy and bit-width.

bit-width assignmentlarge language modelslow-precision floating-point

This work addresses the challenge of balancing dynamic range and precision in numerical representation for memory-constrained embedded and edge AI systems. The authors propose WINT, a weighted integer format that enables flexible trade-offs between precision and range through configurable allocation of mantissa and exponent bits at design time. They derive an analytical model for average relative error and introduce, for the first time, a floating-point-like representation supporting user-defined bit widths. Efficient error evaluation is achieved via O(1)-complexity approximations based on harmonic series and Taylor expansions. Experiments demonstrate that with word lengths of 12 bits or more, WINT using just 2 exponent bits reduces average relative error by 12–33% compared to integer baselines while doubling the dynamic range; with 3 exponent bits, it further extends the range by up to 16× and lowers error by 15–50%.

dynamic rangeedge machine learningembedded systems

Existing methods for predicting distributed training time commonly overlook the impact of floating-point precision—particularly mixed precision—leading to prediction errors as high as 147.85%. This work proposes the first precision-aware training time predictor, which overcomes the limitations of conventional static computation graph approaches by explicitly modeling the dynamic computational and communication overheads under varying precision configurations. By integrating precision-aware modeling, distributed system analysis, and machine learning techniques, the proposed method achieves an average absolute percentage error (MAPE) of 9.8% across diverse precision settings, substantially improving prediction accuracy and robustness in cross-precision scenarios.

distributed trainingfloating-point precisionmixed precision

Hot Scholars

TI

Toshiyuki Imamura

RIKEN Center for Computational Science
Computer ScienceNumerical Linear AlgebraApplied Mathematics
LB

Luca Benini

ETH Zürich, Università di Bologna
Integrated CircuitsComputer ArchitectureEmbedded SystemsVLSI
CX

Cihang Xie

Assistant Professor, University of California, Santa Cruz
Computer VisionMachine Learning
JC

Jianfei Chen

Associate Professor, Tsinghua University
Machine Learning