BASE: Batch-Aware Selection of Experts Using Predicted Removal Error for Efficient MoE Decoding

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the memory bandwidth bottleneck caused by expert set expansion during Mixture-of-Experts (MoE) batch decoding, noting that existing methods lack direct awareness of output errors. This work proposes a batch-aware expert selection mechanism that, for the first time, directly correlates expert retention decisions with output error variations. Specifically, it employs a lightweight linear predictor to estimate the error incurred by expert removal in real time, coupled with custom GPU kernels to accelerate the selection process. Evaluated on Qwen3-30B, the proposed method improves accuracy by 29.5 points. Furthermore, under high computational budgets, it achieves a 60% speedup over dense inference while maintaining an accuracy degradation of less than 0.4 points, demonstrating substantial improvements in both the efficiency and quality of large language model inference.
📝 Abstract
Large language models are increasingly expensive to serve. In large-scale serving systems, autoregressive decoding is often bottlenecked by transferring model weights from accelerator high-bandwidth memory into on-chip SRAM. Mixture-of-experts (MoE) models reduce computation by activating only a small subset of experts per token, but this sparsity does not translate directly to batched decoding. Different requests select different experts; therefore, the combined active set across many concurrent requests can span a substantial fraction of the expert pool and require significantly more expert weights to be transferred. Most expert-reduction techniques make retention decisions independently for each token and therefore do not address this batch-level expansion. More recently, batch-aware methods have attempted to coordinate expert use across concurrent requests and reuse experts already fetched for the batch. Yet their selection criteria are based primarily on router rankings or expert statistics collected during calibration. Consequently, these criteria are not directly tied to the output error caused by dropping an expert, nor do they capture how an expert's contribution changes across tokens at inference time. We instead rank experts according to how much their removal would change the MoE-layer output. To apply this criterion during serving, we train a lightweight linear predictor during calibration that estimates the expert removal cost for each incoming token, and develop custom GPU kernels for cost prediction and expert selection. Across three MoE architectures, BASE improves the quality-efficiency tradeoff without retraining. On Qwen3-30B-A3B, it improves average accuracy by 29.5 points over the strongest baseline at comparable throughput under a tight expert budget. At a higher expert budget, it is 60% faster than dense inference while remaining within 0.4 accuracy points.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
Batched Decoding
Expert Selection
Memory Bandwidth Bottleneck
Large Language Model Serving
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Experts
Batch-Aware Expert Selection
Predicted Removal Error
Efficient Decoding
Custom GPU Kernels