🤖 AI Summary
This study addresses the severe accuracy degradation caused by max-value scaling strategies in low-bit post-training quantization. We theoretically analyze the scaling sensitivity of GPTQ-style methods, proving via probabilistic limit analysis that the normalized loss under Gaussian weights converges to the uniform quantization MSE. This reveals an exponential decay relationship between scaling sensitivity and bit-width, establishing a quantitative link between bit-width and error landscape curvature. Based on these insights, we propose a search-free optimal scaling criterion. Experiments across five large language models demonstrate that, when combined with Hadamard incoherence processing, our method achieves optimal performance at 3 bits or higher without any search procedure, significantly improving the efficiency of low-bit quantization.
📝 Abstract
Post-training quantization (PTQ) methods in the GPTQ family minimize a layer-wise reconstruction error on a uniform grid whose scale must be chosen; the common max-based choice degrades sharply at low bit-widths. We study how sensitive this objective is to the scale. For a layer with i.i.d. Gaussian weights and calibration activations of sufficiently large effective rank, we prove that, as the width grows, the normalized round-to-nearest loss converges with high probability, uniformly over all scales, to the mean-squared error of a uniform quantizer applied to a standard Gaussian; we verify the effective-rank condition for wide, randomly initialized MLPs with odd Lipschitz activations and isotropic Gaussian calibration data. The limiting objective has a unique nondegenerate minimizer, whose scale decreases strictly with the number of levels and whose curvature with respect to relative scale errors decays approximately exponentially with the bit-width. GPTQ experiments on five LLMs show the same trend: the scale rule changes perplexity substantially at 2--3 bits and negligibly from 6 bits on, and a local measure of GPTQ scale sensitivity decreases with bit-width in line with the Gaussian curvature. The Gaussian-optimal scale fails on raw weights; after Hadamard incoherence processing it matches the best searched rule at 3 bits and above without any search, but remains clearly worse at 2 bits.