model quantization

Designs, implements, and evaluates techniques that reduce the numerical precision of machine learning model parameters and activations—converting floating‑point weights/activations to lower‑bit fixed‑point, integer, or binary representations—and integrates these conversions into training and inference pipelines. This includes building post‑training and quantization‑aware training workflows, mixed‑precision schemes, calibration and clipping methods, and hardware‑aware mappings while measuring tradeoffs in accuracy, model size, latency, memory use, and energy.

modelquantization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.93
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$214K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Bitwidth-Specific Logarithmic Arithmetic for Future Hardware-Accelerated Training

Oct 19, 2025
HH
Hassan Hamad
🏛️ University of Southern California

To address the high computational cost and low hardware efficiency of floating-point operations in deep learning training, this paper proposes a hardware-aware low-precision logarithmic fixed-point training method tailored for accelerators. The approach innovatively incorporates a bit-width–aware mechanism into logarithmic addition approximation, jointly optimizing piecewise linear approximation and simulated annealing to achieve Pareto-optimal trade-offs between accuracy and hardware overhead. Leveraging the logarithmic number system (LNS) and bit-accurate C++ simulation, end-to-end training is realized using 12-bit integer arithmetic. Experimental results on VGG-11 and VGG-16 demonstrate accuracy comparable to 32-bit floating-point training, while reducing multiply-accumulate (MAC) unit area by 32.5% and energy consumption by 53.5%. This work establishes a novel hardware–software co-design paradigm for low-precision deep learning training.

Enhancing low-precision logarithmic fixed-point training for hardware acceleratorsOptimizing bitwidth-specific approximations for logarithmic arithmetic operationsReducing area and energy in training with minimal accuracy loss

Standalone 16-bit Neural Network Training: Missing Study for Hardware-Limited Deep Learning Practitioners

May 18, 2023
JY
Juyoung Yun
🏛️ Stony Brook University | State University of New York

Whether 16-bit floating-point (FP16) arithmetic alone can robustly support end-to-end neural network training under hardware resource constraints remains a long-standing, systematically unverified question. Method: This paper introduces the first unified framework for classification-tolerant error analysis, integrating theoretical floating-point error modeling with large-scale empirical validation, and implements fully FP16 forward and backward propagation—without any FP32 “master weights” or loss scaling. Contribution/Results: Rigorous evaluation on CIFAR-10, CIFAR-100, and ImageNet demonstrates that pure FP16 training achieves accuracy parity with FP32 or mixed-precision baselines (accuracy deviation ≤ ±0.1%), while accelerating training by 1.8×. All experiments deploy seamlessly on mainstream GPUs with zero code modification. This work bridges a critical gap between theoretical guarantees and practical feasibility of low-precision training.

Enable efficient neural network training on limited hardwareExplore conditions for 16-bit precision approximationValidate 16-bit precision matches 32-bit accuracy

Algorithms and data structures for automatic precision estimation of neural networks

Sep 29, 2025
IV
Igor V. Netay
🏛️ Joint Stock "Research and production company "Kryptonite" | Institute for Information Transmission Problems, Russian Academy of Sciences

Neural networks suffer from accumulated rounding errors in floating-point arithmetic, causing deviations between actual behavior and mathematical expectations—thereby compromising reliability and interpretability of inference and training. This paper introduces the first automated precision estimation method tailored for deep learning frameworks: it employs lightweight, differentiable data structures and algorithms to enable real-time error propagation tracking during both training and inference, balancing high-fidelity numerical modeling with computational efficiency while seamlessly integrating into mainstream neural network libraries. Its core contribution lies in systematizing and automating floating-point error analysis, enabling end-to-end numerical error monitoring. Extensive experiments across diverse models and tasks demonstrate the method’s broad applicability; they reveal pervasive and significant numerical distortions in most neural networks, underscoring the critical role of precision awareness in ensuring model robustness and trustworthiness.

Addressing computational inaccuracies affecting inference and trainingAutomatically estimating floating-point precision in neural networksEnsuring reliability and interpretability of neural network results

CPT: Efficient Deep Neural Network Training via Cyclic Precision

Jan 25, 2021
YF
Y. Fu
🏛️ Rice University | Facebook Inc

To address low time and energy efficiency in deep neural network training, this paper proposes “cyclic precision scheduling”—a novel paradigm that treats numerical precision as a dynamically optimizable dimension, analogous to learning rate scheduling, by periodically alternating between high and low precision to jointly optimize training. Key contributions include: (i) the first theoretical and empirical demonstration that cyclic precision promotes convergence to wider minima and reduces gradient variance; (ii) an automatic boundary-precision identification mechanism; and (iii) a unified low-precision training framework supporting both floating-point and integer quantization. Experiments across five datasets and eleven models—including image classification and language modeling tasks—show up to 1.8× speedup, 32% energy reduction, no loss in convergence accuracy, an average 0.42% decrease in generalization error, and significantly improved training stability.

Deep Neural NetworksOptimizationTraining Efficiency

Dynamic Range Reduction via Branch-and-Bound

Sep 17, 2024
TG
Thore Gerlach
🏛️ Fraunhofer IAIS

High-precision numerical representations of QUBO coefficients severely degrade the efficiency of specialized hardware—particularly quantum annealers—due to excessive bit-width requirements. Method: We propose a branch-and-bound algorithm that uses dynamic range as a precision-complexity metric, the first such formulation to model dynamic range theory explicitly as the optimization objective for QUBO coefficient quantization. Our framework jointly optimizes precision compression and search efficiency while guaranteeing convergence and solution quality. Crucially, it requires no hardware modification—only input coefficient bit-width reduction via preprocessing. Results: Evaluated on real quantum annealing hardware, our method achieves 35–52% average bit-width compression, 1.8× throughput improvement, 2.3× energy-efficiency gain, and maintains optimal solutions with ≥99.2% fidelity.

Enhancing hardware accelerators via dynamic range reductionOptimizing quantum annealer performance through principled algorithmsReducing precision needs in QUBO problems

Latest Papers

What's happening recently
View more

This work addresses the lack of formal verification foundations for the IEEE P3109 low-precision floating-point standard, whose flexible format and novel features—such as stochastic rounding and saturating arithmetic—pose unique challenges. We present the first complete, parameterized formal model of P3109 in the Lean theorem prover, enabling machine-checkable analysis of its semantics, operations, and key algorithms. Our contributions include the first mechanically verified specification of P3109, a proof that FastTwoSum precisely captures overflow error under saturating arithmetic, and the discovery that ExtractScalar fails at 1-bit precision. The accompanying open-source formal library provides a reusable foundation for the reliable verification of low-precision numerical software.

floating-point arithmeticformal verificationIEEE-P3109

This work addresses the prevalent overuse of double-precision floating-point numbers in numerical programs by proposing an automated mixed-precision tuning methodology. The approach supports user-defined, non-standard low-precision floating-point formats with customizable exponent and mantissa bit-widths, and integrates numerical validation with systematic search within a unified framework to automatically generate program variants that meet prescribed accuracy constraints. Leveraging the PROMISE tool and containerized parallel benchmarking, the method demonstrates that numerous variables across a range of numerical applications and the Rodinia benchmark suite can be safely downgraded in precision. This reduction yields significant improvements in performance while simultaneously decreasing memory consumption and energy usage, all without compromising numerical accuracy.

floating-point formatsmixed-precisionnumerical accuracy

Existing large language models struggle to balance performance and accuracy under uniform low-precision floating-point quantization. This work proposes dMX, the first differentiable mixed-precision quantization framework for the OCP-defined MXFP format. By compressing discrete per-layer bit-width search into a single continuous learnable offset and integrating temperature-annealing scheduling with target-aware regularization, dMX enables a smooth transition from training to hardware deployment. Experiments on Llama, Qwen3, and SmolLM2 demonstrate that dMX consistently outperforms KL-divergence-based heuristic methods across WikiText-2 perplexity and four zero-shot reasoning tasks, achieving Pareto-optimal trade-offs between accuracy and bit-width.

bit-width assignmentlarge language modelslow-precision floating-point

This work proposes a method for training genuine 4-bit convolutional neural networks from scratch on standard CPUs without relying on specialized hardware, custom kernels, or post-training quantization. By integrating tanh-based soft weight clipping, symmetric quantization, dynamic per-layer scaling, and straight-through estimators—all implemented using native PyTorch operations—the approach maintains only 15 unique weight values per layer. It achieves 92.34% and 70.94% accuracy on CIFAR-10 and CIFAR-100, respectively, with less than 0.16% degradation compared to full-precision models, while delivering an 8× memory compression. Notably, the model converges rapidly on mobile devices, reaching 83.16% accuracy within six epochs, marking the first demonstration of near-full-precision performance in 4-bit training on general-purpose CPUs.

4-bit quantizationaccuracy degradationCPU

This study investigates the co-optimization of model scale, dataset size, and numerical precision in low-precision training to balance performance and computational cost. Leveraging a high-dimensional sketching-based linear regression framework, the work models quantization error and analyzes theoretical scaling laws to reveal a fundamental distinction between multiplicative and additive quantization: the former preserves the effective capacity of full-precision models, whereas the latter substantially diminishes it. Theoretical analysis characterizes the intricate coupling among model size, data volume, and precision, and extensive experiments confirm markedly different scaling behaviors between the two quantization paradigms. These findings provide a principled foundation and practical design guidelines for efficient low-precision training.

effective model sizehigh-dimensional linear regressionlow-precision training

Hot Scholars

LB

Luca Benini

ETH Zürich, Università di Bologna
Integrated CircuitsComputer ArchitectureEmbedded SystemsVLSI
MM

Michele Magno

ETH Zurich
Wireless sensor networksSmart Sensors and Internet of ThingsWake up RadioPower management
XY

Xiaofan Yu

University of California Merced
Embedded SystemsEdge Computing
TR

Tajana Rosing

Distinguished Professor, UCSD
computer architecturecyber-physical systemssystem energy efficiency
JL

Junyang Lin

Qwen Team, Alibaba Group & Peking University
Natural Language ProcessingCross-Modal Representation LearningPretraining