BitsMoE: Cost-Aware Bit Allocation in Spectral Space for MoE LLM Quantization

📅 2026-05-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation that existing quantization methods for Mixture-of-Experts (MoE) large language models overlook the structural heterogeneity within linear blocks. To this end, we propose a cost-aware mixed-precision quantization framework. This framework introduces a novel shared-basis spectral decomposition technique to define fine-grained quantization units and combines factorized cost modeling with integer linear programming to achieve component-level optimal bit allocation. Experimental results demonstrate that, under an extremely low 2-bit setting, the proposed method improves accuracy by 2.8% over GEMQ while accelerating quantization speed by 16× and increasing decoding throughput by 6.4×. These findings indicate that our approach effectively balances compression efficiency with inference performance.
📝 Abstract
Mixture-of-Experts (MoE) large language models incur substantial memory costs due to their large expert parameter counts. Mixed-precision quantization reduces these costs by allocating different bit-widths to experts or linear blocks according to their importance. However, assigning a single precision within each expert or linear block overlooks its internal structural heterogeneity. This limitation motivates two key questions: (1) how to define a fine-grained unit for quantization within a linear transformation; and (2) how to characterize the quantization cost of each unit under actual activation patterns and different bit-widths. To address these two questions, we propose BitsMoE, a cost-aware mixed-precision quantization framework built on two complementary techniques: (1) Shared-basis Spectral Decomposition (SSD) separates expert weights into a shared basis and expert-specific spectral components, defining structural quantization units while exploiting cross-expert redundancy. (2) Factorized Quantization Cost Modeling (FQCM) estimates component-wise costs from output reconstruction loss by combining intrinsic spectral importance, activation-dependent importance, and bit-width-dependent distortion. Using these component-wise costs, we formulate bit allocation as an integer linear program (ILP) that minimizes total modeled quantization cost under a fixed memory budget. On Qwen3-30B-A3B at 2-bit, BitsMoE achieves 64.29% average accuracy over seven downstream tasks, outperforming the evaluated MoE-specific methods, including those using ILP-based bit allocation, and exceeding GEMQ by 2.80 percentage points. Under the same setting, it achieves a $16.47\times$ end-to-end offline quantization speedup over GEMQ. It also achieves up to $6.46\times$ the decode throughput of GPTQ.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
LLM Quantization
Mixed-Precision Quantization
Bit Allocation
Structural Heterogeneity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Experts
Mixed-Precision Quantization
Spectral Decomposition
Cost Modeling
Integer Linear Programming
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jiayu Zhao
School of Microelectronics, University of Science and Technology of China
Z
Zihan Teng
School of Microelectronics, University of Science and Technology of China
M
Minhao Fan
College of Computing and Data Science, Nanyang Technological University
Tianrui Ma
Tianrui Ma
Institute of Computing Technology, Chinese Academy of Sciences
visual computingAI hardwareEDAcomputer architecture
W
Wentao Ren
School of Electrical and Electronic Engineering, Nanyang Technological University
Song Chen
Song Chen
University of Science and Technology of China
Physical DesignNetwork-on-ChipsHigh-level SynthesisBrain-Inspired Computing
Weichen Liu
Weichen Liu
College of Computing and Data Science, Nanyang Technological University
Embedded SystemsMultiprocessor SystemsNetwork-on-Chip