quantization-aware training

Designs and implements training procedures that simulate low-bit quantization of neural network weights and intermediate activations during optimization so the resulting model can be converted to few-bit (e.g., ternary or low-bit) representations without large accuracy loss. This work includes building quantization operators and simulated rounding/noise, calibration and lightweight normalization refinements, and fine‑tuning or continual adaptation steps to preserve final-task and reasoning performance after quantization.

quantization-awaretraining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.47
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$229K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

本文针对大模型的内存占用和计算需求问题,通过量化感知训练(QAT)技术,在训练时模拟量化效果,生成低比特模型,保持与全精度模型相近的准确率。

accuracycomputational demandslow-bit models

Must-Read Papers

Most classic and influential ideas
View more

Efficient Multi-bit Quantization Network Training via Weight Bias Correction and Bit-wise Coreset Sampling

Oct 23, 2025
JK
Jinhee Kim
🏛️ Sungkyunkwan University | Korea Advanced Institute of Science and Technology (KAIST)

Existing multi-bit quantized networks require full retraining for each bit-width, leading to linear growth in computational cost with precision levels and necessitating additional fine-tuning for newly introduced precisions. To address this, we propose a single-model, multi-precision training framework. First, we introduce a weight-bias correction technique to align activation distributions across bit-widths and mitigate quantization-induced distribution shifts. Second, we design a gradient-driven, per-bit core-set sampling strategy to enable cross-precision knowledge transfer within a shared backbone. Third, we integrate shared batch normalization and joint training of multiple submodels to support both ResNet and Vision Transformer (ViT) architectures. Our method achieves state-of-the-art accuracy on CIFAR-10/100, TinyImageNet, and ImageNet-1K, while improving training efficiency by up to 7.88× compared to conventional per-precision training—significantly reducing the overhead of deploying multi-precision quantized models.

Addressing linear cost scaling with number of precisionsEliminating fine-tuning needs for multiple precision levelsReducing training overhead in multi-bit quantization networks

Neural network quantization often suffers from significant performance degradation due to parameter discretization. This work proposes a novel quantization paradigm that circumvents the need for straight-through estimators or explicit discretization during training. By exploring low-loss subspaces in the weight space and optimizing a quantization-aware linear path, the method steers quantized models naturally toward high-accuracy regions. Remarkably, directly quantizing only the interpolated points within this subspace achieves accuracy comparable to quantization-aware training and substantially outperforms conventional post-training quantization approaches. The study further highlights the critical role of low-loss subspace structures in enabling efficient and accurate quantization.

Discrete ConstraintsLoss LandscapeLow-Loss Subspaces

Existing neural network quantization methods face two key bottlenecks: non-differentiability causing gradient distortion, and low quantization accuracy—especially for joint weight-activation quantization—along with poor scalability to multi-bit settings. This paper proposes the first fully differentiable bit-shift-based quantization framework, replacing multiplications with bit-shift operations and enabling exact gradient propagation via a differentiable quantization function, with theoretical convergence guarantees. The method supports arbitrary *n*-bit quantization without requiring de novo training; on ResNet-18/ILSVRC-2012, it achieves state-of-the-art 8-bit performance (accuracy drop <1%) after only 15 fine-tuning epochs. During inference, high-precision multiplications are entirely eliminated, replaced solely by lightweight CPU bit-shift instructions—yielding significant reductions in both computational cost and memory footprint.

Achieves near-full-precision accuracy with minimal training epochs and efficient inferenceDevelops differentiable quantization without manual gradient setting in backpropagationEnables multi-bit logarithmic quantization for both weights and activations

Nearly Lossless Adaptive Bit Switching

Feb 03, 2025
HH
Haiduo Huang
🏛️ Xi’an Jiaotong University

To address the challenge of balancing flexibility, efficiency, and accuracy in single-training multi-precision deployment of deep neural networks, this paper proposes a unified multi-precision quantization framework. First, Double Rounding quantization enables near-lossless dynamic switching among INT4/6/8 bit-widths. Second, the Adaptive Learning Rate Scaling (ALRS) mechanism mitigates gradient interference across precision levels. Third, Hessian-Aware Stochastic Bit-switching (HASB) facilitates efficient joint training of mixed-precision models. Our method requires only one quantization-aware training (QAT) pass to support zero-overhead switching across arbitrary integer bit-widths. Evaluated on ImageNet-1K, it significantly outperforms state-of-the-art single-pass multi-precision QAT methods. The framework generalizes effectively to object detection, semantic segmentation, and large language models, reducing storage overhead by up to 75% while incurring negligible accuracy degradation (<0.1% Top-1 drop).

DNN QuantizationFlexibility and EfficiencyPrecision Switching

Extremely low-bit (e.g., 4-bit) post-training quantization (PTQ) suffers severe accuracy degradation due to quantization noise. To address this, we propose the first model-level series expansion framework for PTQ—requiring neither calibration data nor fine-tuning—that losslessly decomposes a floating-point model into multiple low-bit basis models. Our key contributions include: (i) the first application of series expansion to neural network quantization; (ii) a multi-granularity (tensor-, layer-, and model-level) low-bit basis expansion with rigorous convergence guarantees; and (iii) the design of AbelianAdd/Mul operators that form an Abelian group, ensuring parallelizability and commutativity. Experiments demonstrate state-of-the-art performance: 4-bit ResNet-50 achieves 77.03% Top-1 accuracy—surpassing its full-precision baseline—and establishes new SOTA across diverse architectures and tasks in low-bit PTQ.

Enables accurate approximation without calibration or fine-tuningEnsures operation parallelism and commutativity in quantizationReduces performance degradation in low-bit quantization

Latest Papers

What's happening recently
View more

Efficiently Training A Flat Neural Network Before It has been Quantizated

Nov 03, 2025
PX
Peng Xia
🏛️ Beijing University Of Technology

Low-bit post-training quantization (PTQ) of vision transformers suffers from severe accuracy degradation, particularly below 4 bits. Method: This paper proposes Quantization-Aware Flatness Training (QAFT)—a novel PTQ framework that first identifies the flatness of full-precision networks as a critical determinant of low-bit PTQ robustness. By modeling activation and weight quantization errors as independent Gaussian noise, QAFT injects corresponding noise during training and jointly optimizes network parameters to intrinsically adapt the model to the target quantization format. The approach requires no architectural modifications and is model-agnostic. Results: Evaluated on ImageNet, QAFT significantly reduces quantization error across 2–4-bit PTQ regimes. For ViT-B/16, it achieves a +4.2% Top-1 accuracy improvement at 4 bits compared to baseline PTQ methods. QAFT establishes a new, efficient, and broadly applicable paradigm for low-bit quantization-aware training of vision transformers.

Achieving flat neural networks for efficient quantizationReducing quantization error in vision transformersTraining model-agnostic networks for low-bit precision

This work addresses the challenge of deploying deep neural networks on 6G edge devices under extreme compression constraints while preserving model accuracy. Existing mixed-precision quantization methods suffer from coarse granularity, making them ill-suited to capture neuron-level variations in precision requirements. To overcome this limitation, we propose Neuron-level Mixed-Precision Quantization-Aware Training (NMP-QAT), which, for the first time, adaptively assigns discrete bit-widths to individual neurons during training—increasing precision only when necessary—and applies uniformly to both weights and activations. By integrating differentiable proxy functions, straight-through estimators, and a fully discrete inference graph, NMP-QAT significantly outperforms current mixed-precision QAT approaches on MLP and tabular models, achieving superior compression-accuracy trade-offs across both telecom and non-telecom datasets, thereby enabling greener edge AI deployment.

edge AImixed-precision quantizationmodel compression

This work addresses the challenge of efficiently transmitting high-dimensional features from edge devices under stringent constraints on bandwidth, latency, and energy consumption. To this end, the authors propose a trainable, bit-wise soft quantization layer that approximates discrete step functions using multiple sigmoid functions, enabling end-to-end differentiability and task-oriented lossy compression. The method allows users to specify the desired bit-width and can be seamlessly integrated as a lightweight module at the data acquisition stage of neural networks, where it is jointly optimized with downstream tasks. Experimental results across multiple datasets demonstrate that the approach achieves 5–16× compression ratios (relative to 32-bit floating-point representations) using only 2–6 bits per feature, while maintaining accuracy nearly on par with full-precision models—significantly outperforming conventional quantization baselines.

bandwidth constraintedge computingfeature compression

This work addresses the significant performance degradation often observed in low-bit (e.g., 2/4-bit) post-training quantization (PTQ), which stems from the sensitivity of full-precision models to quantization error. To mitigate this issue, the authors propose Efficient Tuning Before Quantization (ETBQ), a lightweight pre-tuning method that injects simulated quantization noise during full-precision training. This guides optimization toward loss landscapes that are inherently robust to quantization, thereby enhancing subsequent PTQ performance. Notably, ETBQ avoids explicit fake quantization operations and yields full-precision models compatible with any PTQ backend without modification. Experimental results demonstrate consistent improvements across benchmarks: under W2A4 settings, it achieves a 2.14% gain in top-1 accuracy on Tiny-ImageNet and a 5.80% increase in mIoU on Cityscapes, validating its effectiveness and broad applicability.

low-bit quantizationperformance degradationpost-training quantization

GDNSQ: Gradual Differentiable Noise Scale Quantization for Low-bit Neural Networks

Aug 19, 2025
SS
Sergey Salishev
🏛️ aifoundry.org | Ainekko Co.

To address the capacity degradation and difficulty in dynamically tracking bottlenecks in low-bit quantized neural networks—caused by non-differentiable rounding operations—this paper models quantization as a cascade of noisy channels and proposes a progressively differentiable noise-scaled quantization framework. Methodologically, it jointly optimizes learnable bit-widths, noise scales, and clipping ranges via straight-through estimators (STE) for end-to-end training; incorporates an outlier-penalty term to precisely enforce target bit-widths; and integrates lightweight knowledge distillation to enhance training stability and gradient smoothness. The approach achieves competitive accuracy under the extremely challenging W1A1 setting—surpassing state-of-the-art methods significantly—while maintaining efficient forward inference. Key contributions include: (i) a noise-modeling-driven differentiable quantization paradigm, and (ii) a multi-variable co-constraint mechanism unifying bit-width, noise, and clipping optimization.

Fine-tuning addresses quantization bottlenecks via constrained optimizationMethod enables low-bit W1A1 networks while maintaining STE efficiencyQuantized neural networks face capacity reduction from rounding noise

Hot Scholars

HQ

Haotong Qin

ETH Zürich
TinyMLModel CompressionComputer VisionDeep Learning
DA

Dan Alistarh

Professor at IST Austria
Machine LearningAlgorithmsDistributed Computing
MM

Michele Magno

ETH Zurich
Wireless sensor networksSmart Sensors and Internet of ThingsWake up RadioPower management
MS

Muhammad Shafique

Professor, ECE, New York University (AD-UAE, Tandon-USA), Director eBRAIN Lab
Embedded Machine LearningBrain-Inspired ComputingRobust & Energy-Efficient System DesignSmart
LK

Linghe Kong

Shanghai Jiao Tong University
Internet of ThingsMobile computingBig data