Score
Designs, implements, or evaluates algorithms, tools and deployment pipelines that reduce the numerical precision of trained model parameters and, optionally, activations after training to enable lower-memory, lower-bandwidth, or lower-latency inference. This work encompasses data-agnostic/data-free and dynamic methods, weight-only and weight‑activation schemes, ternary/4-bit/mixed‑precision/product and per‑tile quantization variants, hybrid approaches, and the measurements and analyses of accuracy, quantization noise, collapse points and runtime/storage tradeoffs.
Deploying deep neural networks (DNNs) faces challenges from high computational overhead and large model sizes. While low-bit weight quantization accelerates inference and reduces memory bandwidth requirements, it often incurs substantial accuracy degradation. This paper presents a systematic survey of low-bit weight quantization research from 2019 to 2024. We propose the first unified taxonomy comprising eight major categories and 24 subcategories—covering linear/nonlinear quantization, layer-wise/channel-wise calibration, retraining-free and fine-tuning-based paradigms, gradient approximation techniques, and mixed-precision search strategies. Through structured comparative analysis of over 100 state-of-the-art works, we identify common bottlenecks, clarify promising future directions, and highlight open challenges. To foster reproducibility and industrial adoption, we open-source Awesome-Model-Quantization—a curated, continuously updated resource repository—thereby advancing standardization and practical deployment of quantization techniques.
To address high latency and excessive memory consumption in large language model (LLM) inference, this paper proposes W4A8—a hardware-friendly dual-precision quantization paradigm: weights are stored as 4-bit integers, while activations and computations use 8-bit floating-point (FP8) arithmetic, jointly optimizing storage efficiency and computational performance without compromising accuracy. We introduce Dual-Precision Quantization (DPQ), an overhead-free quantization algorithm that employs hardware-aware post-training calibration to eliminate the latency overhead typically incurred by conventional accuracy-compensation techniques. DPQ requires no hardware modifications and is compatible with diverse modern AI accelerators. Experimental results demonstrate up to 35–62% higher inference throughput, 75% reduction in weight memory footprint, and less than 0.5% accuracy degradation relative to FP16 baselines—achieving up to 1.8× speedup.
Whether 16-bit floating-point (FP16) arithmetic alone can robustly support end-to-end neural network training under hardware resource constraints remains a long-standing, systematically unverified question. Method: This paper introduces the first unified framework for classification-tolerant error analysis, integrating theoretical floating-point error modeling with large-scale empirical validation, and implements fully FP16 forward and backward propagation—without any FP32 “master weights” or loss scaling. Contribution/Results: Rigorous evaluation on CIFAR-10, CIFAR-100, and ImageNet demonstrates that pure FP16 training achieves accuracy parity with FP32 or mixed-precision baselines (accuracy deviation ≤ ±0.1%), while accelerating training by 1.8×. All experiments deploy seamlessly on mainstream GPUs with zero code modification. This work bridges a critical gap between theoretical guarantees and practical feasibility of low-precision training.
To address the prohibitively high memory overhead of stateful optimizers (e.g., Adam) in large language model training—whose auxiliary states can reach up to twice the parameter count—this work proposes the first stable 2-bit ultra-low-precision optimizer. Our method introduces: (1) a logarithmic quantization scheme robust to signal drowning, and (2) an EMA-based dynamic modeling mechanism guided by gradient variance analysis, coupled with momentum self-adaptation to ensure convergence stability under extreme quantization. Evaluated on a 7B-parameter model, our optimizer reduces optimizer-state memory by approximately 45 GB while incurring negligible accuracy degradation. This advance substantially improves the feasibility of large-model training under severe memory constraints, enabling efficient optimization with unprecedented bit-width reduction.
Existing low-precision quantization methods for sub-microsecond neural network inference on FPGAs often incur substantial accuracy degradation, while conventional mixed-precision approaches suffer from coarse granularity and inflexible optimization. Method: This paper proposes a gradient-based fine-grained mixed-precision quantization method. It is the first to model bitwidth as a learnable parameter embedded within a quantization-aware training framework, enabling layer-wise independent and differentiable bitwidth assignment for weights and activations. The method jointly optimizes bitwidth configurations under FPGA/ASIC hardware constraints. Contribution/Results: Experiments demonstrate that, without sacrificing model accuracy, the approach reduces hardware resource consumption to 5% of the baseline and compresses end-to-end inference latency to 20%—achieving a fivefold speedup. This significantly enhances deployment efficiency of on-chip neural networks for ultra-low-latency applications.
Edge NPU cross-platform low-bit deployment suffers from opaque and heterogeneous compiler quantization strategies—such as scaling, clipping, and kernel support—leading to significant accuracy fluctuations for the same floating-point model across backends and necessitating repeated manual tuning. To address this, we propose Quant-Trim: a training-time method that jointly integrates progressive pseudo-quantization with backward pruning to suppress outlier-induced scale inflation, thereby enhancing model robustness against diverse hardware quantization schemes. Quant-Trim is hardware-agnostic and supports symmetric/asymmetric, per-tensor/per-channel, and INT8/INT4 configurations without computational graph modification or vendor-specific calibration. Experiments demonstrate that it substantially narrows the accuracy gap between floating-point and low-bit models, reduces reliance on compiler-level tuning, and achieves consistent high performance across platforms in key edge inference metrics—including latency, throughput, energy consumption, and cost.
Today, large language models have demonstrated their strengths in various tasks ranging from reasoning, code generation, and complex problem solving. However, this advancement comes with a high computational cost and memory requirements, making it challenging to deploy these models on edge devices to ensure real-time responses and data privacy. Quantization is one common approach to reducing memory use, but most methods apply it uniformly across all layers. This does not account for the fact that different layers may respond differently to reduced precision. Importantly, memory consumption and computational throughput are not necessarily aligned, further complicating deployment decisions. This paper proposes an adaptive mixed precision quantization mechanism that balances memory, latency, and accuracy in edge deployment under user-defined priorities. This is achieved by analyzing the layer-wise contribution and by inferring how different quantization types behave across the target hardware platform in order to assign the most suitable quantization type to each layer. This integration ensures that layer importance and the overall performance trade-offs are jointly respected in this design. Our work unlocks new configuration designs that uniform quantization cannot achieve, expanding the solution space to efficiently deploy the LLMs on resource-constrained devices.
This work addresses the quantization error in deep neural networks caused by diverse data distributions by proposing a unified framework grounded in statistical error analysis, which for the first time tightly integrates quantizer design with the characteristics of data distributions. The framework encompasses an iterative optimization-based quantizer applicable to arbitrary distributions and an analytical quantizer tailored for near-Gaussian weight distributions, supporting both integer and floating-point formats and seamlessly integrating with quantization-aware training (QAT). Experimental results demonstrate that the proposed approach significantly improves accuracy and training stability of low-bit models across various architectures and datasets, outperforming existing uniform and floating-point quantization methods.
Neural networks suffer from accumulated rounding errors in floating-point arithmetic, causing deviations between actual behavior and mathematical expectations—thereby compromising reliability and interpretability of inference and training. This paper introduces the first automated precision estimation method tailored for deep learning frameworks: it employs lightweight, differentiable data structures and algorithms to enable real-time error propagation tracking during both training and inference, balancing high-fidelity numerical modeling with computational efficiency while seamlessly integrating into mainstream neural network libraries. Its core contribution lies in systematizing and automating floating-point error analysis, enabling end-to-end numerical error monitoring. Extensive experiments across diverse models and tasks demonstrate the method’s broad applicability; they reveal pervasive and significant numerical distortions in most neural networks, underscoring the critical role of precision awareness in ensuring model robustness and trustworthiness.
This work addresses the lack of systematic, multi-granularity comparison between low-precision floating-point (FP) and integer (INT) quantization for handling LLM activation outliers on AI hardware (e.g., Blackwell architecture). We conduct algorithm–hardware co-design, proposing three techniques: fine-grained block-wise quantization, Hadamard rotation to suppress outliers, and symmetric clipping to eliminate gradient bias in low-bit training. We首次 demonstrate that MXINT8 achieves near-lossless training at 4 bits—surpassing MXFP4/NVFP4 in both accuracy and energy efficiency. Moreover, NVINT4, when combined with outlier control, outperforms NVFP4. These findings challenge the FP-dominant paradigm in AI hardware design, establishing a unified evaluation framework for low-bit quantization and introducing an integer-first, co-optimized path for efficient LLM inference and training.