post-training quantization

Designs, implements, or evaluates algorithms, tools and deployment pipelines that reduce the numerical precision of trained model parameters and, optionally, activations after training to enable lower-memory, lower-bandwidth, or lower-latency inference. This work encompasses data-agnostic/data-free and dynamic methods, weight-only and weight‑activation schemes, ternary/4-bit/mixed‑precision/product and per‑tile quantization variants, hybrid approaches, and the measurements and analyses of accuracy, quantization noise, collapse points and runtime/storage tradeoffs.

post-trainingquantization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.26
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$228K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

To address high latency and excessive memory consumption in large language model (LLM) inference, this paper proposes W4A8—a hardware-friendly dual-precision quantization paradigm: weights are stored as 4-bit integers, while activations and computations use 8-bit floating-point (FP8) arithmetic, jointly optimizing storage efficiency and computational performance without compromising accuracy. We introduce Dual-Precision Quantization (DPQ), an overhead-free quantization algorithm that employs hardware-aware post-training calibration to eliminate the latency overhead typically incurred by conventional accuracy-compensation techniques. DPQ requires no hardware modifications and is compatible with diverse modern AI accelerators. Experimental results demonstrate up to 35–62% higher inference throughput, 75% reduction in weight memory footprint, and less than 0.5% accuracy degradation relative to FP16 baselines—achieving up to 1.8× speedup.

Enhancing hardware utilization with mixed-precision computationMinimizing accuracy loss during low-bit quantizationReducing model size and latency for efficient DNN inference

Standalone 16-bit Neural Network Training: Missing Study for Hardware-Limited Deep Learning Practitioners

May 18, 2023
JY
Juyoung Yun
🏛️ Stony Brook University | State University of New York

Whether 16-bit floating-point (FP16) arithmetic alone can robustly support end-to-end neural network training under hardware resource constraints remains a long-standing, systematically unverified question. Method: This paper introduces the first unified framework for classification-tolerant error analysis, integrating theoretical floating-point error modeling with large-scale empirical validation, and implements fully FP16 forward and backward propagation—without any FP32 “master weights” or loss scaling. Contribution/Results: Rigorous evaluation on CIFAR-10, CIFAR-100, and ImageNet demonstrates that pure FP16 training achieves accuracy parity with FP32 or mixed-precision baselines (accuracy deviation ≤ ±0.1%), while accelerating training by 1.8×. All experiments deploy seamlessly on mainstream GPUs with zero code modification. This work bridges a critical gap between theoretical guarantees and practical feasibility of low-precision training.

Enable efficient neural network training on limited hardwareExplore conditions for 16-bit precision approximationValidate 16-bit precision matches 32-bit accuracy

Pushing the Limits of Low-Bit Optimizers: A Focus on EMA Dynamics

May 01, 2025
CX
Cong Xu
🏛️ East China Normal University | Tencent | WeChat AI | Tsinghua University | Shanghai Innovation Institute

To address the prohibitively high memory overhead of stateful optimizers (e.g., Adam) in large language model training—whose auxiliary states can reach up to twice the parameter count—this work proposes the first stable 2-bit ultra-low-precision optimizer. Our method introduces: (1) a logarithmic quantization scheme robust to signal drowning, and (2) an EMA-based dynamic modeling mechanism guided by gradient variance analysis, coupled with momentum self-adaptation to ensure convergence stability under extreme quantization. Evaluated on a 7B-parameter model, our optimizer reduces optimizer-state memory by approximately 45 GB while incurring negligible accuracy degradation. This advance substantially improves the feasibility of large-model training under severe memory constraints, enabling efficient optimization with unprecedented bit-width reduction.

Addressing signal swamping in unsigned 3-2 bit quantizationMitigating gradient variance in signed low-bit optimizer statesReducing optimizer memory overhead via ultra-low-precision quantization

Existing low-precision quantization methods for sub-microsecond neural network inference on FPGAs often incur substantial accuracy degradation, while conventional mixed-precision approaches suffer from coarse granularity and inflexible optimization. Method: This paper proposes a gradient-based fine-grained mixed-precision quantization method. It is the first to model bitwidth as a learnable parameter embedded within a quantization-aware training framework, enabling layer-wise independent and differentiable bitwidth assignment for weights and activations. The method jointly optimizes bitwidth configurations under FPGA/ASIC hardware constraints. Contribution/Results: Experiments demonstrate that, without sacrificing model accuracy, the approach reduces hardware resource consumption to 5% of the baseline and compresses end-to-end inference latency to 20%—achieving a fivefold speedup. This significantly enhances deployment efficiency of on-chip neural networks for ultra-low-latency applications.

Enables sub-microsecond inference latency on FPGA hardwareOptimizes neural network parameter bit-widths via gradient descentReduces resource consumption while maintaining model accuracy

Latest Papers

What's happening recently
View more

Edge NPU cross-platform low-bit deployment suffers from opaque and heterogeneous compiler quantization strategies—such as scaling, clipping, and kernel support—leading to significant accuracy fluctuations for the same floating-point model across backends and necessitating repeated manual tuning. To address this, we propose Quant-Trim: a training-time method that jointly integrates progressive pseudo-quantization with backward pruning to suppress outlier-induced scale inflation, thereby enhancing model robustness against diverse hardware quantization schemes. Quant-Trim is hardware-agnostic and supports symmetric/asymmetric, per-tensor/per-channel, and INT8/INT4 configurations without computational graph modification or vendor-specific calibration. Experiments demonstrate that it substantially narrows the accuracy gap between floating-point and low-bit models, reduces reliance on compiler-level tuning, and achieves consistent high performance across platforms in key edge inference metrics—including latency, throughput, energy consumption, and cost.

Addresses inconsistent accuracy across edge NPUs due to vendor compiler variationsEliminates need for per-backend model retraining and vendor-specific modificationsReduces accuracy gap between floating-point and low-bit quantization deployments

Today, large language models have demonstrated their strengths in various tasks ranging from reasoning, code generation, and complex problem solving. However, this advancement comes with a high computational cost and memory requirements, making it challenging to deploy these models on edge devices to ensure real-time responses and data privacy. Quantization is one common approach to reducing memory use, but most methods apply it uniformly across all layers. This does not account for the fact that different layers may respond differently to reduced precision. Importantly, memory consumption and computational throughput are not necessarily aligned, further complicating deployment decisions. This paper proposes an adaptive mixed precision quantization mechanism that balances memory, latency, and accuracy in edge deployment under user-defined priorities. This is achieved by analyzing the layer-wise contribution and by inferring how different quantization types behave across the target hardware platform in order to assign the most suitable quantization type to each layer. This integration ensures that layer importance and the overall performance trade-offs are jointly respected in this design. Our work unlocks new configuration designs that uniform quantization cannot achieve, expanding the solution space to efficiently deploy the LLMs on resource-constrained devices.

edge deploymentlarge language modelslayer-wise sensitivity

This work addresses the quantization error in deep neural networks caused by diverse data distributions by proposing a unified framework grounded in statistical error analysis, which for the first time tightly integrates quantizer design with the characteristics of data distributions. The framework encompasses an iterative optimization-based quantizer applicable to arbitrary distributions and an analytical quantizer tailored for near-Gaussian weight distributions, supporting both integer and floating-point formats and seamlessly integrating with quantization-aware training (QAT). Experimental results demonstrate that the proposed approach significantly improves accuracy and training stability of low-bit models across various architectures and datasets, outperforming existing uniform and floating-point quantization methods.

data distributionsdeep neural networkslow-precision inference

Algorithms and data structures for automatic precision estimation of neural networks

Sep 29, 2025
IV
Igor V. Netay
🏛️ Joint Stock "Research and production company "Kryptonite" | Institute for Information Transmission Problems, Russian Academy of Sciences

Neural networks suffer from accumulated rounding errors in floating-point arithmetic, causing deviations between actual behavior and mathematical expectations—thereby compromising reliability and interpretability of inference and training. This paper introduces the first automated precision estimation method tailored for deep learning frameworks: it employs lightweight, differentiable data structures and algorithms to enable real-time error propagation tracking during both training and inference, balancing high-fidelity numerical modeling with computational efficiency while seamlessly integrating into mainstream neural network libraries. Its core contribution lies in systematizing and automating floating-point error analysis, enabling end-to-end numerical error monitoring. Extensive experiments across diverse models and tasks demonstrate the method’s broad applicability; they reveal pervasive and significant numerical distortions in most neural networks, underscoring the critical role of precision awareness in ensuring model robustness and trustworthiness.

Addressing computational inaccuracies affecting inference and trainingAutomatically estimating floating-point precision in neural networksEnsuring reliability and interpretability of neural network results

INT v.s. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization Formats

Oct 29, 2025
MC
Mengzhao Chen
🏛️ The University of Hong Kong | PicoHeart | ByteDance Seed

This work addresses the lack of systematic, multi-granularity comparison between low-precision floating-point (FP) and integer (INT) quantization for handling LLM activation outliers on AI hardware (e.g., Blackwell architecture). We conduct algorithm–hardware co-design, proposing three techniques: fine-grained block-wise quantization, Hadamard rotation to suppress outliers, and symmetric clipping to eliminate gradient bias in low-bit training. We首次 demonstrate that MXINT8 achieves near-lossless training at 4 bits—surpassing MXFP4/NVFP4 in both accuracy and energy efficiency. Moreover, NVINT4, when combined with outlier control, outperforms NVFP4. These findings challenge the FP-dominant paradigm in AI hardware design, establishing a unified evaluation framework for low-bit quantization and introducing an integer-first, co-optimized path for efficient LLM inference and training.

Analyzes performance crossover between coarse-grained FP and fine-grained INT formatsChallenges one-size-fits-all FP approach for future AI acceleratorsCompares FP and INT quantization formats across granularities for AI hardware

Hot Scholars

HQ

Haotong Qin

ETH Zürich
TinyMLModel CompressionComputer VisionDeep Learning
MM

Michele Magno

ETH Zurich
Wireless sensor networksSmart Sensors and Internet of ThingsWake up RadioPower management
DA

Dan Alistarh

Professor at IST Austria
Machine LearningAlgorithmsDistributed Computing
JC

Jianfei Chen

Associate Professor, Tsinghua University
Machine Learning
QG

Qingyi Gu

Institute of Automation, Chinese Academy of Sciences
High-speed visioncell analysis