Score
Designs, implements, and evaluates numerical precision schemes for training and inference of machine learning models in which weights, activations, optimizer moments, and compute kernels use different floating- or integer-bit widths (e.g., bf16 weights with fp32 moments) to reduce runtime and energy while maintaining model accuracy. This includes sensitivity-guided and inter-/intra-layer bit-allocation algorithms, mixed-precision programming and low-precision kernels, precision-aware optimizers and validation workflows (benchmarking throughput vs accuracy, aggregating metrics such as perplexity and reasoning signals), and techniques to prevent numerical cascades and preserve stability.
Whether 16-bit floating-point (FP16) arithmetic alone can robustly support end-to-end neural network training under hardware resource constraints remains a long-standing, systematically unverified question. Method: This paper introduces the first unified framework for classification-tolerant error analysis, integrating theoretical floating-point error modeling with large-scale empirical validation, and implements fully FP16 forward and backward propagation—without any FP32 “master weights” or loss scaling. Contribution/Results: Rigorous evaluation on CIFAR-10, CIFAR-100, and ImageNet demonstrates that pure FP16 training achieves accuracy parity with FP32 or mixed-precision baselines (accuracy deviation ≤ ±0.1%), while accelerating training by 1.8×. All experiments deploy seamlessly on mainstream GPUs with zero code modification. This work bridges a critical gap between theoretical guarantees and practical feasibility of low-precision training.
Neural networks suffer from accumulated rounding errors in floating-point arithmetic, causing deviations between actual behavior and mathematical expectations—thereby compromising reliability and interpretability of inference and training. This paper introduces the first automated precision estimation method tailored for deep learning frameworks: it employs lightweight, differentiable data structures and algorithms to enable real-time error propagation tracking during both training and inference, balancing high-fidelity numerical modeling with computational efficiency while seamlessly integrating into mainstream neural network libraries. Its core contribution lies in systematizing and automating floating-point error analysis, enabling end-to-end numerical error monitoring. Extensive experiments across diverse models and tasks demonstrate the method’s broad applicability; they reveal pervasive and significant numerical distortions in most neural networks, underscoring the critical role of precision awareness in ensuring model robustness and trustworthiness.
To address the high computational cost and low hardware efficiency of floating-point operations in deep learning training, this paper proposes a hardware-aware low-precision logarithmic fixed-point training method tailored for accelerators. The approach innovatively incorporates a bit-width–aware mechanism into logarithmic addition approximation, jointly optimizing piecewise linear approximation and simulated annealing to achieve Pareto-optimal trade-offs between accuracy and hardware overhead. Leveraging the logarithmic number system (LNS) and bit-accurate C++ simulation, end-to-end training is realized using 12-bit integer arithmetic. Experimental results on VGG-11 and VGG-16 demonstrate accuracy comparable to 32-bit floating-point training, while reducing multiply-accumulate (MAC) unit area by 32.5% and energy consumption by 53.5%. This work establishes a novel hardware–software co-design paradigm for low-precision deep learning training.
This work addresses the reliance of stochastic rounding (SR) on high-entropy random bits in low-precision (e.g., FP16) and mixed-precision computing, revealing a previously overlooked systematic bias induced by finite-bit random sources (FBSRs)—a bias invisible under infinite-precision theory yet detrimental to numerical reliability. Through rigorous error modeling, floating-point rounding analysis, and empirical training (e.g., ResNet-18), we quantitatively characterize the bias magnitude across multiple FBSR schemes for the first time, demonstrating up to a 1.2% degradation in training accuracy. Our study extends the reliability assessment framework for low-precision computation by explicitly incorporating FBSR-induced bias as a critical dimension. We further propose a low-bit SR implementation framework that explicitly controls bias while preserving efficiency, and release open-source, reproducible code. This work bridges theoretical SR analysis and practical low-precision system design, enabling more robust and predictable stochastic quantization in deep learning accelerators.
Conventional FP32 multipliers incur substantial hardware overhead and poor energy efficiency, hindering high-performance CNN inference. Method: This paper proposes a co-optimization framework for deploying error-bounded approximate FP32 multipliers *within* convolution kernels, targeting CNN inference. Leveraging the IEEE 754 standard, it employs approximate compressors for significand multiplication and integrates NSGA-II—a multi-objective genetic algorithm—to jointly optimize multiplier type selection, placement, and composition order, thereby balancing accuracy and hardware efficiency. Contribution/Results: Evaluated across multiple CNN models, the approach maintains >99% of the original accuracy while reducing multiplier area and power consumption by 32.7% and 28.4% on average, respectively. This yields significant improvements in inference energy efficiency and throughput, providing a deployable, architecture-level solution for high-precision approximate computing.
This work addresses the prevalent overuse of double-precision floating-point numbers in numerical programs by proposing an automated mixed-precision tuning methodology. The approach supports user-defined, non-standard low-precision floating-point formats with customizable exponent and mantissa bit-widths, and integrates numerical validation with systematic search within a unified framework to automatically generate program variants that meet prescribed accuracy constraints. Leveraging the PROMISE tool and containerized parallel benchmarking, the method demonstrates that numerous variables across a range of numerical applications and the Rodinia benchmark suite can be safely downgraded in precision. This reduction yields significant improvements in performance while simultaneously decreasing memory consumption and energy usage, all without compromising numerical accuracy.
Existing large language models struggle to balance performance and accuracy under uniform low-precision floating-point quantization. This work proposes dMX, the first differentiable mixed-precision quantization framework for the OCP-defined MXFP format. By compressing discrete per-layer bit-width search into a single continuous learnable offset and integrating temperature-annealing scheduling with target-aware regularization, dMX enables a smooth transition from training to hardware deployment. Experiments on Llama, Qwen3, and SmolLM2 demonstrate that dMX consistently outperforms KL-divergence-based heuristic methods across WikiText-2 perplexity and four zero-shot reasoning tasks, achieving Pareto-optimal trade-offs between accuracy and bit-width.
This work addresses the challenge of balancing dynamic range and precision in numerical representation for memory-constrained embedded and edge AI systems. The authors propose WINT, a weighted integer format that enables flexible trade-offs between precision and range through configurable allocation of mantissa and exponent bits at design time. They derive an analytical model for average relative error and introduce, for the first time, a floating-point-like representation supporting user-defined bit widths. Efficient error evaluation is achieved via O(1)-complexity approximations based on harmonic series and Taylor expansions. Experiments demonstrate that with word lengths of 12 bits or more, WINT using just 2 exponent bits reduces average relative error by 12–33% compared to integer baselines while doubling the dynamic range; with 3 exponent bits, it further extends the range by up to 16× and lowers error by 15–50%.
Existing methods for predicting distributed training time commonly overlook the impact of floating-point precision—particularly mixed precision—leading to prediction errors as high as 147.85%. This work proposes the first precision-aware training time predictor, which overcomes the limitations of conventional static computation graph approaches by explicitly modeling the dynamic computational and communication overheads under varying precision configurations. By integrating precision-aware modeling, distributed system analysis, and machine learning techniques, the proposed method achieves an average absolute percentage error (MAPE) of 9.8% across diverse precision settings, substantially improving prediction accuracy and robustness in cross-precision scenarios.