🤖 AI Summary
This study addresses the inefficiency of existing low-bit quantization methods for large language models (LLMs), where uneven codeword allocation leads to poor codebook utilization and storage bottlenecks. To overcome this, we propose BARQ, a novel framework that introduces entropy-regularized optimal transport to achieve balanced soft assignment. The resulting optimization problem is efficiently solved via the Sinkhorn algorithm. Furthermore, BARQ refines the codebook by integrating a curvature-weighted reconstruction loss with an assignment-weighted centroid update strategy, for which we establish theoretical optimality under fixed assignments. Extensive experiments across mainstream LLMs demonstrate that BARQ significantly reduces perplexity and improves zero-shot accuracy, consistently outperforming existing baseline methods. This work offers a promising new paradigm for efficient large model quantization.
📝 Abstract
As large language models (LLMs) grow in parameter count, model storage and parameter memory traffic have become major bottlenecks to efficient deployment. Codebook-based weight quantization reduces these costs, but imbalanced nearest-codeword assignments during fitting can leave some codewords insufficiently updated, limiting effective codebook utilization. To address this limitation, we propose Balanced Assignment Refinement for Quantization (BARQ), which improves quantization quality through balanced fitting of existing codebooks. Specifically, we first compute joint soft assignments between weight blocks and codewords through entropically regularized optimal transport with uniform marginals and curvature-weighted reconstruction costs, ensuring equal positive fitting mass for every codeword in the exact solution. We then refine the codewords through an assignment-weighted barycentric update, which we prove minimizes the fitting objective for fixed assignments. For finite Sinkhorn iterations, the implemented update retains this optimality provided all codeword masses exceed the denominator floor. Finally, we discard the soft assignments and use the refined codebook for standard hard nearest-codeword encoding, with our analysis establishing sufficient conditions for reducing hard-quantization distortion and evaluation loss. Across multiple LLMs, BARQ achieves lower perplexity and higher mean zero-shot accuracy than the evaluated baselines at comparable bit budgets. The code is available at https://github.com/chenhangcuisg-code/BARQ.