Score
Designs, implements, and evaluates techniques that reduce the numerical precision of machine learning model parameters and activations—converting floating‑point weights/activations to lower‑bit fixed‑point, integer, or binary representations—and integrates these conversions into training and inference pipelines. This includes building post‑training and quantization‑aware training workflows, mixed‑precision schemes, calibration and clipping methods, and hardware‑aware mappings while measuring tradeoffs in accuracy, model size, latency, memory use, and energy.
Deploying deep neural networks (DNNs) faces challenges from high computational overhead and large model sizes. While low-bit weight quantization accelerates inference and reduces memory bandwidth requirements, it often incurs substantial accuracy degradation. This paper presents a systematic survey of low-bit weight quantization research from 2019 to 2024. We propose the first unified taxonomy comprising eight major categories and 24 subcategories—covering linear/nonlinear quantization, layer-wise/channel-wise calibration, retraining-free and fine-tuning-based paradigms, gradient approximation techniques, and mixed-precision search strategies. Through structured comparative analysis of over 100 state-of-the-art works, we identify common bottlenecks, clarify promising future directions, and highlight open challenges. To foster reproducibility and industrial adoption, we open-source Awesome-Model-Quantization—a curated, continuously updated resource repository—thereby advancing standardization and practical deployment of quantization techniques.
To address the high computational cost and low hardware efficiency of floating-point operations in deep learning training, this paper proposes a hardware-aware low-precision logarithmic fixed-point training method tailored for accelerators. The approach innovatively incorporates a bit-width–aware mechanism into logarithmic addition approximation, jointly optimizing piecewise linear approximation and simulated annealing to achieve Pareto-optimal trade-offs between accuracy and hardware overhead. Leveraging the logarithmic number system (LNS) and bit-accurate C++ simulation, end-to-end training is realized using 12-bit integer arithmetic. Experimental results on VGG-11 and VGG-16 demonstrate accuracy comparable to 32-bit floating-point training, while reducing multiply-accumulate (MAC) unit area by 32.5% and energy consumption by 53.5%. This work establishes a novel hardware–software co-design paradigm for low-precision deep learning training.
Whether 16-bit floating-point (FP16) arithmetic alone can robustly support end-to-end neural network training under hardware resource constraints remains a long-standing, systematically unverified question. Method: This paper introduces the first unified framework for classification-tolerant error analysis, integrating theoretical floating-point error modeling with large-scale empirical validation, and implements fully FP16 forward and backward propagation—without any FP32 “master weights” or loss scaling. Contribution/Results: Rigorous evaluation on CIFAR-10, CIFAR-100, and ImageNet demonstrates that pure FP16 training achieves accuracy parity with FP32 or mixed-precision baselines (accuracy deviation ≤ ±0.1%), while accelerating training by 1.8×. All experiments deploy seamlessly on mainstream GPUs with zero code modification. This work bridges a critical gap between theoretical guarantees and practical feasibility of low-precision training.
Neural networks suffer from accumulated rounding errors in floating-point arithmetic, causing deviations between actual behavior and mathematical expectations—thereby compromising reliability and interpretability of inference and training. This paper introduces the first automated precision estimation method tailored for deep learning frameworks: it employs lightweight, differentiable data structures and algorithms to enable real-time error propagation tracking during both training and inference, balancing high-fidelity numerical modeling with computational efficiency while seamlessly integrating into mainstream neural network libraries. Its core contribution lies in systematizing and automating floating-point error analysis, enabling end-to-end numerical error monitoring. Extensive experiments across diverse models and tasks demonstrate the method’s broad applicability; they reveal pervasive and significant numerical distortions in most neural networks, underscoring the critical role of precision awareness in ensuring model robustness and trustworthiness.
To address low time and energy efficiency in deep neural network training, this paper proposes “cyclic precision scheduling”—a novel paradigm that treats numerical precision as a dynamically optimizable dimension, analogous to learning rate scheduling, by periodically alternating between high and low precision to jointly optimize training. Key contributions include: (i) the first theoretical and empirical demonstration that cyclic precision promotes convergence to wider minima and reduces gradient variance; (ii) an automatic boundary-precision identification mechanism; and (iii) a unified low-precision training framework supporting both floating-point and integer quantization. Experiments across five datasets and eleven models—including image classification and language modeling tasks—show up to 1.8× speedup, 32% energy reduction, no loss in convergence accuracy, an average 0.42% decrease in generalization error, and significantly improved training stability.
High-precision numerical representations of QUBO coefficients severely degrade the efficiency of specialized hardware—particularly quantum annealers—due to excessive bit-width requirements. Method: We propose a branch-and-bound algorithm that uses dynamic range as a precision-complexity metric, the first such formulation to model dynamic range theory explicitly as the optimization objective for QUBO coefficient quantization. Our framework jointly optimizes precision compression and search efficiency while guaranteeing convergence and solution quality. Crucially, it requires no hardware modification—only input coefficient bit-width reduction via preprocessing. Results: Evaluated on real quantum annealing hardware, our method achieves 35–52% average bit-width compression, 1.8× throughput improvement, 2.3× energy-efficiency gain, and maintains optimal solutions with ≥99.2% fidelity.
This work addresses the lack of formal verification foundations for the IEEE P3109 low-precision floating-point standard, whose flexible format and novel features—such as stochastic rounding and saturating arithmetic—pose unique challenges. We present the first complete, parameterized formal model of P3109 in the Lean theorem prover, enabling machine-checkable analysis of its semantics, operations, and key algorithms. Our contributions include the first mechanically verified specification of P3109, a proof that FastTwoSum precisely captures overflow error under saturating arithmetic, and the discovery that ExtractScalar fails at 1-bit precision. The accompanying open-source formal library provides a reusable foundation for the reliable verification of low-precision numerical software.
This work addresses the prevalent overuse of double-precision floating-point numbers in numerical programs by proposing an automated mixed-precision tuning methodology. The approach supports user-defined, non-standard low-precision floating-point formats with customizable exponent and mantissa bit-widths, and integrates numerical validation with systematic search within a unified framework to automatically generate program variants that meet prescribed accuracy constraints. Leveraging the PROMISE tool and containerized parallel benchmarking, the method demonstrates that numerous variables across a range of numerical applications and the Rodinia benchmark suite can be safely downgraded in precision. This reduction yields significant improvements in performance while simultaneously decreasing memory consumption and energy usage, all without compromising numerical accuracy.
Existing large language models struggle to balance performance and accuracy under uniform low-precision floating-point quantization. This work proposes dMX, the first differentiable mixed-precision quantization framework for the OCP-defined MXFP format. By compressing discrete per-layer bit-width search into a single continuous learnable offset and integrating temperature-annealing scheduling with target-aware regularization, dMX enables a smooth transition from training to hardware deployment. Experiments on Llama, Qwen3, and SmolLM2 demonstrate that dMX consistently outperforms KL-divergence-based heuristic methods across WikiText-2 perplexity and four zero-shot reasoning tasks, achieving Pareto-optimal trade-offs between accuracy and bit-width.
This work proposes a method for training genuine 4-bit convolutional neural networks from scratch on standard CPUs without relying on specialized hardware, custom kernels, or post-training quantization. By integrating tanh-based soft weight clipping, symmetric quantization, dynamic per-layer scaling, and straight-through estimators—all implemented using native PyTorch operations—the approach maintains only 15 unique weight values per layer. It achieves 92.34% and 70.94% accuracy on CIFAR-10 and CIFAR-100, respectively, with less than 0.16% degradation compared to full-precision models, while delivering an 8× memory compression. Notably, the model converges rapidly on mobile devices, reaching 83.16% accuracy within six epochs, marking the first demonstration of near-full-precision performance in 4-bit training on general-purpose CPUs.
This study investigates the co-optimization of model scale, dataset size, and numerical precision in low-precision training to balance performance and computational cost. Leveraging a high-dimensional sketching-based linear regression framework, the work models quantization error and analyzes theoretical scaling laws to reveal a fundamental distinction between multiplicative and additive quantization: the former preserves the effective capacity of full-precision models, whereas the latter substantially diminishes it. Theoretical analysis characterizes the intricate coupling among model size, data volume, and precision, and extensive experiments confirm markedly different scaling behaviors between the two quantization paradigms. These findings provide a principled foundation and practical design guidelines for efficient low-precision training.