per-channel int4 calibration

Designs and implements per-channel 4-bit quantization and calibration procedures that compute and apply independent scale (and optional zero-point) parameters per tensor channel to convert weights and/or activations to int4 format. Builds conversion and inference pipelines and analyzes quantization error accumulation, memory/compute reductions, and accuracy tradeoffs to restore or preserve global reasoning accuracy (often without retraining).

per-channelint4calibration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.26
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study challenges the common assumption that convergence in FP32 implies suitability for quantization by revealing a sharp degradation in INT4 performance after FP32 convergence. Analyzing all 154 training checkpoints of Pythia-160M using calibration-free per-group INT4 probing, the authors identify a three-phase explosive growth in quantization error post-convergence and, for the first time, link the onset of this collapse to FP32 perplexity convergence rather than learning rate decay. They propose Oscillatory Lock-In, a novel learning rate schedule combined with kurtosis-based outlier filtering, which significantly enhances INT4 robustness. In multi-schedule comparisons, this approach effectively mitigates the surge in late-stage quantization error—from 11% to 517%—reducing it on average by 2.2 percentage points (p<0.0001), thereby validating the critical role of schedule amplitude calibration.

FP32 convergenceINT4 quantizationpost-training quantization

Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling

Dec 01, 2025
JC
Jack Cook
🏛️ Massachusetts Institute of Technology | NVIDIA

To address forward/backward divergence and degraded inference performance caused by NVFP4 low-precision quantization in large language model (LLM) training, this paper proposes the Four Over Six (4/6) quantization method. The core innovation lies in dynamically evaluating two candidate scaling factors per data block, prioritizing representation accuracy near the maximum absolute value region to substantially mitigate performance degradation dominated by FP4’s inherent quantization error. Built upon adaptive block-wise floating-point scaling, 4/6 is compatible with both forward and backward passes, as well as multiple post-training quantization paradigms, and features an optimized implementation for NVIDIA Blackwell architectures. Experiments demonstrate stable convergence during Transformer- and hybrid-architecture pretraining, with training loss closely matching BF16 baselines and consistent improvements in downstream task accuracy.

Addresses NVFP4 quantization errors causing training divergence and inference degradation.Enables efficient NVFP4 training on GPUs while enhancing model accuracy.Introduces adaptive block scaling to improve representation of near-maximal values.

This study addresses the severe accuracy degradation caused by max-value scaling strategies in low-bit post-training quantization. We theoretically analyze the scaling sensitivity of GPTQ-style methods, proving via probabilistic limit analysis that the normalized loss under Gaussian weights converges to the uniform quantization MSE. This reveals an exponential decay relationship between scaling sensitivity and bit-width, establishing a quantitative link between bit-width and error landscape curvature. Based on these insights, we propose a search-free optimal scaling criterion. Experiments across five large language models demonstrate that, when combined with Hadamard incoherence processing, our method achieves optimal performance at 3 bits or higher without any search procedure, significantly improving the efficiency of low-bit quantization.

GPTQLow-Bit QuantizationPost-Training Quantization

This work addresses the severe degradation in reasoning performance of large language models under 4-bit quantization, particularly the loss of accuracy in low-entropy symbols such as digits and operators, which existing post-training quantization (PTQ) and quantization-aware training (QAT) methods struggle to recover. The authors propose ReQAT, a novel framework that identifies low-entropy tokens during inference as quantization-sensitive points and introduces three core techniques: Trace-Aligned QAT, Selective Entropy Minimization, and Quantization-Friendly Initialization (Q-FIT), collectively optimizing critical decision points. Combined with a RoPE-consistent KV cache transformation and enhancements to the FP4 format, ReQAT achieves higher accuracy than BF16 fine-tuning under full W4A4KV4 quantization and delivers up to 3.9× and 3.1× throughput speedups on NVIDIA DGX Spark and B200 systems, respectively, within the same training budget.

4-bit quantizationKV cache quantizationlow-entropy tokens

Cat: Post-training quantization error reduction via cluster-based affine transformation

Sep 30, 2025
AZ
Ali Zoljodi
🏛️ Mälardalen University | University of Wurzburg

Low-bit post-training quantization (PTQ) often incurs substantial accuracy degradation, exacerbated by conventional uniform affine transformations. To address this, we propose Clustering-based Affine Transformation (CAT), a method that learns cluster-specific affine parameters for distinct output clusters—thereby aligning quantized and full-precision output distributions with near-zero parameter overhead. CAT operates as a plug-and-play module requiring no fine-tuning or retraining, enabling seamless integration into existing PTQ pipelines. On ImageNet-1K, CAT achieves 53.18% Top-1 accuracy for W2A2 ResNet-18—surpassing the state-of-the-art by over 3%. It demonstrates consistent robustness across diverse architectures and quantization configurations. The core innovation lies in coupling clustering analysis with cluster-level affine calibration, effectively mitigating distribution mismatch—a critical challenge in low-bit PTQ.

Enhancing quantization performance without fine-tuning parametersImproving output alignment between quantized and full-precision modelsReducing accuracy loss in low-bit post-training quantization

Latest Papers

What's happening recently
View more

Existing post-training quantization (PTQ) methods suffer significant performance degradation on complex tasks such as mathematical reasoning and code generation due to their neglect of the reasoning dynamics inherent in large language models. To address this, this work proposes ScaleQ-1.58, a novel framework that, for the first time, incorporates the model’s own generated reasoning trajectories into PTQ calibration through an “Attend Your Own Thoughts” (AYOT) strategy. Combined with the differentiable ternarization method CAT-Q, ScaleQ-1.58 achieves highly efficient 1.58-bit quantization using only 4M calibration tokens. The approach attains 90.52% of the performance of the BitNet b1.58 2B4T baseline on Qwen3-1.7B and yields an absolute improvement of 8.97% on Qwen3-4B, while reducing calibration overhead by six orders of magnitude. It further demonstrates exceptional generalization and scalability across models up to 235B parameters and diverse reasoning tasks.

calibrationchain-of-thought reasoningpost-training quantization

This work presents the first application of MixQ in conjunction with SmoothQuant to enable efficient 4-bit (HiF4/MXFP4) inference for the Wan2.2-I2V-A14B image-to-video large model. To address heavy-tailed activation distributions, the authors propose a dual-branch quantization architecture that combines channel-wise smoothing to compress dynamic ranges with block-wise HiF4 packing and dual-branch GEMM. After calibration, outlier columns are retained in higher precision while the remaining channels undergo strict W4A4 quantization. Evaluated on VBench I2V, the method achieves performance within only 2–3.5% of FP16 across most metrics, substantially outperforming the native HiFloat4 baseline—which degrades by approximately 5%—and notably enhances motion smoothness in generated videos.

activation outlierslarge modellow-bit inference

This study addresses the combinatorial complexity of bit-width allocation and the sensitivity to calibration data in mixed-precision quantization. To tackle these challenges, we propose a post-training mixed-precision quantization method based on hierarchical probabilistic error attribution. By constructing a separable scoring mechanism and conducting probabilistic local perturbation analysis, the proposed approach achieves efficient and robust bit-width allocation without relying on external solvers. Experimental results demonstrate that our method accelerates the allocation process by up to 2570× and yields a PSNR gain of 7.5 dB. Furthermore, it significantly outperforms existing baseline methods in both computational efficiency and resilience against data corruption, establishing a highly effective solution for practical mixed-precision quantization scenarios.

bit allocationcorrupted calibration datamixed-precision quantization

This work addresses the challenge of obtaining optimal scaling factors in post-training quantization, where conventional data-free heuristics often fall short. The authors propose PiSO, an algorithm that, for the first time, enables precise and efficient optimization of channel-wise (and grouped) scaling factors under round-to-nearest quantization. By partitioning the search space into a finite set of intervals and deriving closed-form optimal solutions within each interval—augmented with an error correction strategy—PiSO significantly enhances low-bit quantization performance. Extensive experiments on Llama and Qwen model families demonstrate consistent improvements across varying model scales and bit widths, with notable reductions in perplexity and gains in zero-shot accuracy, particularly pronounced in ultra-low-bit regimes.

large language modelslow-bit representationpost-training quantization

This study addresses the reliance of existing sensitivity estimation methods on extensive calibration data, which hinders precise mixed-precision allocation under strict budgets in large model quantization. By identifying the spectral flatness property of quantization errors, this work proposes a data-free sensitivity estimator requiring only a single random Gaussian probe. It quantifies inter-layer sensitivity via Frobenius norm and propagation-based scoring, subsequently employing a knapsack algorithm to achieve optimal bit allocation under budget constraints. This approach overcomes traditional data-dependency bottlenecks, enabling accurate mixed-precision quantization with minimal computational overhead across diverse architectures. Experimental results demonstrate that the proposed method significantly reduces perplexity, outperforming both uniform 4-bit quantization and existing baseline approaches.

budget-targeteddata-freemixed-precision quantization

Hot Scholars

HQ

Haotong Qin

ETH Zürich
TinyMLModel CompressionComputer VisionDeep Learning
MM

Michele Magno

ETH Zurich
Wireless sensor networksSmart Sensors and Internet of ThingsWake up RadioPower management
WL

WeiHsiang Liao

Sony Research Inc.
Signal ProcessingMachine Learning
HL

Haizhou Li

The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen), China; NUS, Singapore
Automatic Speech RecognitionSpeaker RecognitionLanguage RecognitionVoice Conversion