🤖 AI Summary
This study addresses the limitation that existing quantization methods for Mixture-of-Experts (MoE) large language models overlook the structural heterogeneity within linear blocks. To this end, we propose a cost-aware mixed-precision quantization framework. This framework introduces a novel shared-basis spectral decomposition technique to define fine-grained quantization units and combines factorized cost modeling with integer linear programming to achieve component-level optimal bit allocation. Experimental results demonstrate that, under an extremely low 2-bit setting, the proposed method improves accuracy by 2.8% over GEMQ while accelerating quantization speed by 16× and increasing decoding throughput by 6.4×. These findings indicate that our approach effectively balances compression efficiency with inference performance.
📝 Abstract
Mixture-of-Experts (MoE) large language models incur substantial memory costs due to their large expert parameter counts. Mixed-precision quantization reduces these costs by allocating different bit-widths to experts or linear blocks according to their importance. However, assigning a single precision within each expert or linear block overlooks its internal structural heterogeneity. This limitation motivates two key questions: (1) how to define a fine-grained unit for quantization within a linear transformation; and (2) how to characterize the quantization cost of each unit under actual activation patterns and different bit-widths. To address these two questions, we propose BitsMoE, a cost-aware mixed-precision quantization framework built on two complementary techniques: (1) Shared-basis Spectral Decomposition (SSD) separates expert weights into a shared basis and expert-specific spectral components, defining structural quantization units while exploiting cross-expert redundancy. (2) Factorized Quantization Cost Modeling (FQCM) estimates component-wise costs from output reconstruction loss by combining intrinsic spectral importance, activation-dependent importance, and bit-width-dependent distortion. Using these component-wise costs, we formulate bit allocation as an integer linear program (ILP) that minimizes total modeled quantization cost under a fixed memory budget. On Qwen3-30B-A3B at 2-bit, BitsMoE achieves 64.29% average accuracy over seven downstream tasks, outperforming the evaluated MoE-specific methods, including those using ILP-based bit allocation, and exceeding GEMQ by 2.80 percentage points. Under the same setting, it achieves a $16.47\times$ end-to-end offline quantization speedup over GEMQ. It also achieves up to $6.46\times$ the decode throughput of GPTQ.