🤖 AI Summary
This study addresses the accuracy degradation and error coupling in group quantization of large language models caused by the misalignment between quantization grids and integer codes. We formulate quantization as a bilinear mixed-integer least-squares problem and propose a plug-and-play joint alternating refinement method. By leveraging bounded Babai proposals and joint least-squares fitting, our approach alternately optimizes group scales and codes to enable synergistic multi-code updates. Notably, this method requires no backpropagation and introduces zero inference overhead. Extensive experiments demonstrate that our approach significantly reduces perplexity in 90 out of 96 comparisons across mainstream models, achieving up to a 36% perplexity reduction for 3-bit Round-To-Nearest (RTN) quantization while completing the process in under one minute per 7B parameters.
📝 Abstract
Group-wise post-training quantizers for large language models round weights onto a grid that is not refit to the resulting integer codes. We show that this leaves accuracy on the table: the best grid depends on the codes, input correlations couple the errors of different groups, and useful code changes often involve many codes at once. We propose JARQ , a plug-in refinement that starts from any group-wise quantizer and alternates a joint least-squares fit of all group scales with bounded Babai proposals that move many codes of a group together on the current grid. The problem is a bilinear box-constrained mixed-integer least-squares problem; the solver is backpropagation-free, does not increase the layer-wise objective under exact scale solves, and keeps the host's bit width, groups, zero points, and inference cost. Across Llama-2, Llama-3, and Qwen models with RTN, GPTQ, OmniQuant, and AWQ hosts, JARQ lowers perplexity in 90 of 96 comparisons, cuts three-bit RTN perplexity by up to 36%, raises mean multiple-choice accuracy in 23 of 24 configurations, and improves QEP, QuaRot, and OJBKQ outputs, at under a minute per 7B block.