Score
Designs and implements training procedures that simulate low-bit quantization of neural network weights and intermediate activations during optimization so the resulting model can be converted to few-bit (e.g., ternary or low-bit) representations without large accuracy loss. This work includes building quantization operators and simulated rounding/noise, calibration and lightweight normalization refinements, and fine‑tuning or continual adaptation steps to preserve final-task and reasoning performance after quantization.
Deploying deep neural networks (DNNs) faces challenges from high computational overhead and large model sizes. While low-bit weight quantization accelerates inference and reduces memory bandwidth requirements, it often incurs substantial accuracy degradation. This paper presents a systematic survey of low-bit weight quantization research from 2019 to 2024. We propose the first unified taxonomy comprising eight major categories and 24 subcategories—covering linear/nonlinear quantization, layer-wise/channel-wise calibration, retraining-free and fine-tuning-based paradigms, gradient approximation techniques, and mixed-precision search strategies. Through structured comparative analysis of over 100 state-of-the-art works, we identify common bottlenecks, clarify promising future directions, and highlight open challenges. To foster reproducibility and industrial adoption, we open-source Awesome-Model-Quantization—a curated, continuously updated resource repository—thereby advancing standardization and practical deployment of quantization techniques.
本文针对大模型的内存占用和计算需求问题,通过量化感知训练(QAT)技术,在训练时模拟量化效果,生成低比特模型,保持与全精度模型相近的准确率。
Existing multi-bit quantized networks require full retraining for each bit-width, leading to linear growth in computational cost with precision levels and necessitating additional fine-tuning for newly introduced precisions. To address this, we propose a single-model, multi-precision training framework. First, we introduce a weight-bias correction technique to align activation distributions across bit-widths and mitigate quantization-induced distribution shifts. Second, we design a gradient-driven, per-bit core-set sampling strategy to enable cross-precision knowledge transfer within a shared backbone. Third, we integrate shared batch normalization and joint training of multiple submodels to support both ResNet and Vision Transformer (ViT) architectures. Our method achieves state-of-the-art accuracy on CIFAR-10/100, TinyImageNet, and ImageNet-1K, while improving training efficiency by up to 7.88× compared to conventional per-precision training—significantly reducing the overhead of deploying multi-precision quantized models.
Neural network quantization often suffers from significant performance degradation due to parameter discretization. This work proposes a novel quantization paradigm that circumvents the need for straight-through estimators or explicit discretization during training. By exploring low-loss subspaces in the weight space and optimizing a quantization-aware linear path, the method steers quantized models naturally toward high-accuracy regions. Remarkably, directly quantizing only the interpolated points within this subspace achieves accuracy comparable to quantization-aware training and substantially outperforms conventional post-training quantization approaches. The study further highlights the critical role of low-loss subspace structures in enabling efficient and accurate quantization.
Existing neural network quantization methods face two key bottlenecks: non-differentiability causing gradient distortion, and low quantization accuracy—especially for joint weight-activation quantization—along with poor scalability to multi-bit settings. This paper proposes the first fully differentiable bit-shift-based quantization framework, replacing multiplications with bit-shift operations and enabling exact gradient propagation via a differentiable quantization function, with theoretical convergence guarantees. The method supports arbitrary *n*-bit quantization without requiring de novo training; on ResNet-18/ILSVRC-2012, it achieves state-of-the-art 8-bit performance (accuracy drop <1%) after only 15 fine-tuning epochs. During inference, high-precision multiplications are entirely eliminated, replaced solely by lightweight CPU bit-shift instructions—yielding significant reductions in both computational cost and memory footprint.
To address the challenge of balancing flexibility, efficiency, and accuracy in single-training multi-precision deployment of deep neural networks, this paper proposes a unified multi-precision quantization framework. First, Double Rounding quantization enables near-lossless dynamic switching among INT4/6/8 bit-widths. Second, the Adaptive Learning Rate Scaling (ALRS) mechanism mitigates gradient interference across precision levels. Third, Hessian-Aware Stochastic Bit-switching (HASB) facilitates efficient joint training of mixed-precision models. Our method requires only one quantization-aware training (QAT) pass to support zero-overhead switching across arbitrary integer bit-widths. Evaluated on ImageNet-1K, it significantly outperforms state-of-the-art single-pass multi-precision QAT methods. The framework generalizes effectively to object detection, semantic segmentation, and large language models, reducing storage overhead by up to 75% while incurring negligible accuracy degradation (<0.1% Top-1 drop).
Extremely low-bit (e.g., 4-bit) post-training quantization (PTQ) suffers severe accuracy degradation due to quantization noise. To address this, we propose the first model-level series expansion framework for PTQ—requiring neither calibration data nor fine-tuning—that losslessly decomposes a floating-point model into multiple low-bit basis models. Our key contributions include: (i) the first application of series expansion to neural network quantization; (ii) a multi-granularity (tensor-, layer-, and model-level) low-bit basis expansion with rigorous convergence guarantees; and (iii) the design of AbelianAdd/Mul operators that form an Abelian group, ensuring parallelizability and commutativity. Experiments demonstrate state-of-the-art performance: 4-bit ResNet-50 achieves 77.03% Top-1 accuracy—surpassing its full-precision baseline—and establishes new SOTA across diverse architectures and tasks in low-bit PTQ.
Low-bit post-training quantization (PTQ) of vision transformers suffers from severe accuracy degradation, particularly below 4 bits. Method: This paper proposes Quantization-Aware Flatness Training (QAFT)—a novel PTQ framework that first identifies the flatness of full-precision networks as a critical determinant of low-bit PTQ robustness. By modeling activation and weight quantization errors as independent Gaussian noise, QAFT injects corresponding noise during training and jointly optimizes network parameters to intrinsically adapt the model to the target quantization format. The approach requires no architectural modifications and is model-agnostic. Results: Evaluated on ImageNet, QAFT significantly reduces quantization error across 2–4-bit PTQ regimes. For ViT-B/16, it achieves a +4.2% Top-1 accuracy improvement at 4 bits compared to baseline PTQ methods. QAFT establishes a new, efficient, and broadly applicable paradigm for low-bit quantization-aware training of vision transformers.
This work addresses the challenge of deploying deep neural networks on 6G edge devices under extreme compression constraints while preserving model accuracy. Existing mixed-precision quantization methods suffer from coarse granularity, making them ill-suited to capture neuron-level variations in precision requirements. To overcome this limitation, we propose Neuron-level Mixed-Precision Quantization-Aware Training (NMP-QAT), which, for the first time, adaptively assigns discrete bit-widths to individual neurons during training—increasing precision only when necessary—and applies uniformly to both weights and activations. By integrating differentiable proxy functions, straight-through estimators, and a fully discrete inference graph, NMP-QAT significantly outperforms current mixed-precision QAT approaches on MLP and tabular models, achieving superior compression-accuracy trade-offs across both telecom and non-telecom datasets, thereby enabling greener edge AI deployment.
This work addresses the challenge of efficiently transmitting high-dimensional features from edge devices under stringent constraints on bandwidth, latency, and energy consumption. To this end, the authors propose a trainable, bit-wise soft quantization layer that approximates discrete step functions using multiple sigmoid functions, enabling end-to-end differentiability and task-oriented lossy compression. The method allows users to specify the desired bit-width and can be seamlessly integrated as a lightweight module at the data acquisition stage of neural networks, where it is jointly optimized with downstream tasks. Experimental results across multiple datasets demonstrate that the approach achieves 5–16× compression ratios (relative to 32-bit floating-point representations) using only 2–6 bits per feature, while maintaining accuracy nearly on par with full-precision models—significantly outperforming conventional quantization baselines.
This work addresses the significant performance degradation often observed in low-bit (e.g., 2/4-bit) post-training quantization (PTQ), which stems from the sensitivity of full-precision models to quantization error. To mitigate this issue, the authors propose Efficient Tuning Before Quantization (ETBQ), a lightweight pre-tuning method that injects simulated quantization noise during full-precision training. This guides optimization toward loss landscapes that are inherently robust to quantization, thereby enhancing subsequent PTQ performance. Notably, ETBQ avoids explicit fake quantization operations and yields full-precision models compatible with any PTQ backend without modification. Experimental results demonstrate consistent improvements across benchmarks: under W2A4 settings, it achieves a 2.14% gain in top-1 accuracy on Tiny-ImageNet and a 5.80% increase in mIoU on Cityscapes, validating its effectiveness and broad applicability.
To address the capacity degradation and difficulty in dynamically tracking bottlenecks in low-bit quantized neural networks—caused by non-differentiable rounding operations—this paper models quantization as a cascade of noisy channels and proposes a progressively differentiable noise-scaled quantization framework. Methodologically, it jointly optimizes learnable bit-widths, noise scales, and clipping ranges via straight-through estimators (STE) for end-to-end training; incorporates an outlier-penalty term to precisely enforce target bit-widths; and integrates lightweight knowledge distillation to enhance training stability and gradient smoothness. The approach achieves competitive accuracy under the extremely challenging W1A1 setting—surpassing state-of-the-art methods significantly—while maintaining efficient forward inference. Key contributions include: (i) a noise-modeling-driven differentiable quantization paradigm, and (ii) a multi-variable co-constraint mechanism unifying bit-width, noise, and clipping optimization.