Score
Designs and implements per-channel 4-bit quantization and calibration procedures that compute and apply independent scale (and optional zero-point) parameters per tensor channel to convert weights and/or activations to int4 format. Builds conversion and inference pipelines and analyzes quantization error accumulation, memory/compute reductions, and accuracy tradeoffs to restore or preserve global reasoning accuracy (often without retraining).
Deploying deep neural networks (DNNs) faces challenges from high computational overhead and large model sizes. While low-bit weight quantization accelerates inference and reduces memory bandwidth requirements, it often incurs substantial accuracy degradation. This paper presents a systematic survey of low-bit weight quantization research from 2019 to 2024. We propose the first unified taxonomy comprising eight major categories and 24 subcategories—covering linear/nonlinear quantization, layer-wise/channel-wise calibration, retraining-free and fine-tuning-based paradigms, gradient approximation techniques, and mixed-precision search strategies. Through structured comparative analysis of over 100 state-of-the-art works, we identify common bottlenecks, clarify promising future directions, and highlight open challenges. To foster reproducibility and industrial adoption, we open-source Awesome-Model-Quantization—a curated, continuously updated resource repository—thereby advancing standardization and practical deployment of quantization techniques.
This study challenges the common assumption that convergence in FP32 implies suitability for quantization by revealing a sharp degradation in INT4 performance after FP32 convergence. Analyzing all 154 training checkpoints of Pythia-160M using calibration-free per-group INT4 probing, the authors identify a three-phase explosive growth in quantization error post-convergence and, for the first time, link the onset of this collapse to FP32 perplexity convergence rather than learning rate decay. They propose Oscillatory Lock-In, a novel learning rate schedule combined with kurtosis-based outlier filtering, which significantly enhances INT4 robustness. In multi-schedule comparisons, this approach effectively mitigates the surge in late-stage quantization error—from 11% to 517%—reducing it on average by 2.2 percentage points (p<0.0001), thereby validating the critical role of schedule amplitude calibration.
To address forward/backward divergence and degraded inference performance caused by NVFP4 low-precision quantization in large language model (LLM) training, this paper proposes the Four Over Six (4/6) quantization method. The core innovation lies in dynamically evaluating two candidate scaling factors per data block, prioritizing representation accuracy near the maximum absolute value region to substantially mitigate performance degradation dominated by FP4’s inherent quantization error. Built upon adaptive block-wise floating-point scaling, 4/6 is compatible with both forward and backward passes, as well as multiple post-training quantization paradigms, and features an optimized implementation for NVIDIA Blackwell architectures. Experiments demonstrate stable convergence during Transformer- and hybrid-architecture pretraining, with training loss closely matching BF16 baselines and consistent improvements in downstream task accuracy.
This study addresses the severe accuracy degradation caused by max-value scaling strategies in low-bit post-training quantization. We theoretically analyze the scaling sensitivity of GPTQ-style methods, proving via probabilistic limit analysis that the normalized loss under Gaussian weights converges to the uniform quantization MSE. This reveals an exponential decay relationship between scaling sensitivity and bit-width, establishing a quantitative link between bit-width and error landscape curvature. Based on these insights, we propose a search-free optimal scaling criterion. Experiments across five large language models demonstrate that, when combined with Hadamard incoherence processing, our method achieves optimal performance at 3 bits or higher without any search procedure, significantly improving the efficiency of low-bit quantization.
This work addresses the severe degradation in reasoning performance of large language models under 4-bit quantization, particularly the loss of accuracy in low-entropy symbols such as digits and operators, which existing post-training quantization (PTQ) and quantization-aware training (QAT) methods struggle to recover. The authors propose ReQAT, a novel framework that identifies low-entropy tokens during inference as quantization-sensitive points and introduces three core techniques: Trace-Aligned QAT, Selective Entropy Minimization, and Quantization-Friendly Initialization (Q-FIT), collectively optimizing critical decision points. Combined with a RoPE-consistent KV cache transformation and enhancements to the FP4 format, ReQAT achieves higher accuracy than BF16 fine-tuning under full W4A4KV4 quantization and delivers up to 3.9× and 3.1× throughput speedups on NVIDIA DGX Spark and B200 systems, respectively, within the same training budget.
Low-bit post-training quantization (PTQ) often incurs substantial accuracy degradation, exacerbated by conventional uniform affine transformations. To address this, we propose Clustering-based Affine Transformation (CAT), a method that learns cluster-specific affine parameters for distinct output clusters—thereby aligning quantized and full-precision output distributions with near-zero parameter overhead. CAT operates as a plug-and-play module requiring no fine-tuning or retraining, enabling seamless integration into existing PTQ pipelines. On ImageNet-1K, CAT achieves 53.18% Top-1 accuracy for W2A2 ResNet-18—surpassing the state-of-the-art by over 3%. It demonstrates consistent robustness across diverse architectures and quantization configurations. The core innovation lies in coupling clustering analysis with cluster-level affine calibration, effectively mitigating distribution mismatch—a critical challenge in low-bit PTQ.
Existing post-training quantization (PTQ) methods suffer significant performance degradation on complex tasks such as mathematical reasoning and code generation due to their neglect of the reasoning dynamics inherent in large language models. To address this, this work proposes ScaleQ-1.58, a novel framework that, for the first time, incorporates the model’s own generated reasoning trajectories into PTQ calibration through an “Attend Your Own Thoughts” (AYOT) strategy. Combined with the differentiable ternarization method CAT-Q, ScaleQ-1.58 achieves highly efficient 1.58-bit quantization using only 4M calibration tokens. The approach attains 90.52% of the performance of the BitNet b1.58 2B4T baseline on Qwen3-1.7B and yields an absolute improvement of 8.97% on Qwen3-4B, while reducing calibration overhead by six orders of magnitude. It further demonstrates exceptional generalization and scalability across models up to 235B parameters and diverse reasoning tasks.
This work presents the first application of MixQ in conjunction with SmoothQuant to enable efficient 4-bit (HiF4/MXFP4) inference for the Wan2.2-I2V-A14B image-to-video large model. To address heavy-tailed activation distributions, the authors propose a dual-branch quantization architecture that combines channel-wise smoothing to compress dynamic ranges with block-wise HiF4 packing and dual-branch GEMM. After calibration, outlier columns are retained in higher precision while the remaining channels undergo strict W4A4 quantization. Evaluated on VBench I2V, the method achieves performance within only 2–3.5% of FP16 across most metrics, substantially outperforming the native HiFloat4 baseline—which degrades by approximately 5%—and notably enhances motion smoothness in generated videos.
This study addresses the combinatorial complexity of bit-width allocation and the sensitivity to calibration data in mixed-precision quantization. To tackle these challenges, we propose a post-training mixed-precision quantization method based on hierarchical probabilistic error attribution. By constructing a separable scoring mechanism and conducting probabilistic local perturbation analysis, the proposed approach achieves efficient and robust bit-width allocation without relying on external solvers. Experimental results demonstrate that our method accelerates the allocation process by up to 2570× and yields a PSNR gain of 7.5 dB. Furthermore, it significantly outperforms existing baseline methods in both computational efficiency and resilience against data corruption, establishing a highly effective solution for practical mixed-precision quantization scenarios.
This work addresses the challenge of obtaining optimal scaling factors in post-training quantization, where conventional data-free heuristics often fall short. The authors propose PiSO, an algorithm that, for the first time, enables precise and efficient optimization of channel-wise (and grouped) scaling factors under round-to-nearest quantization. By partitioning the search space into a finite set of intervals and deriving closed-form optimal solutions within each interval—augmented with an error correction strategy—PiSO significantly enhances low-bit quantization performance. Extensive experiments on Llama and Qwen model families demonstrate consistent improvements across varying model scales and bit widths, with notable reductions in perplexity and gains in zero-shot accuracy, particularly pronounced in ultra-low-bit regimes.
This study addresses the reliance of existing sensitivity estimation methods on extensive calibration data, which hinders precise mixed-precision allocation under strict budgets in large model quantization. By identifying the spectral flatness property of quantization errors, this work proposes a data-free sensitivity estimator requiring only a single random Gaussian probe. It quantifies inter-layer sensitivity via Frobenius norm and propagation-based scoring, subsequently employing a knapsack algorithm to achieve optimal bit allocation under budget constraints. This approach overcomes traditional data-dependency bottlenecks, enabling accurate mixed-precision quantization with minimal computational overhead across diverse architectures. Experimental results demonstrate that the proposed method significantly reduces perplexity, outperforming both uniform 4-bit quantization and existing baseline approaches.