SlimWise: Decoupling Expert Pruning Across Prefill and Decode for Efficient MoE Serving

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the expert weight transmission bottleneck during the decoding phase of Mixture-of-Experts (MoE) models and the degradation of prefill quality caused by conventional pruning. We propose a disaggregated inference framework that employs the full model for prefilling while reusing a pruned model for decoding with direct KV cache inheritance. The core innovation lies in introducing the first training-free cross-phase KV cache handoff mechanism, which decouples expert pool configurations across the two stages and incorporates lightweight distillation to compensate for accuracy loss. Implemented on vLLM, the system supports both prefill-decode disaggregation and co-located deployment. Evaluated on the Qwen3.6 model, our approach achieves a 1.81× improvement in decoding throughput at a 50% expert pruning ratio with negligible accuracy degradation.
📝 Abstract
Mixture-of-experts (MoE) models activate few experts per token, yet batched decoding can access nearly the entire expert pool, making expert-weight traffic a major bottleneck. Expert pruning reduces this traffic, but conventional approaches also prune compute-bound prefill, sacrificing model quality for little throughput benefit. We present SlimWise, a serving framework that tailors the expert pool to each inference phase. SlimWise performs prefill with the full model and decode with a pruned model that directly reuses the prefill-generated KV cache without conversion. Across two MoE backbones and three pruning criteria, this training-free KV cache handoff substantially narrows accuracy gaps relative to the full model in many settings. We also show that benchmark accuracy can conceal substantial pruning-induced changes in generation length. To address these distortions and residual accuracy loss, SlimWise introduces a low-cost distillation stage that trains the decoder to continue from full-model KV caches while updating only a small subset of parameters. Implemented in vLLM, SlimWise supports both prefill-decode (PD) disaggregation and PD-colocated serving. On Qwen3.6-35B-A3B, SlimWise improves decode throughput by up to 1.81x at 50% expert pruning with minimal accuracy loss.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
expert pruning
MoE serving
prefill-decode
KV cache
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Experts
Expert Pruning
KV Cache Handoff
Knowledge Distillation
LLM Serving
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Gunho Park
Gunho Park
NAVER Cloud
K
Kyoungho Jeun
a2sys
J
Juntaek Oh
a2sys
B
Byeongjun Shin
a2sys, KAIST
B
Baeseong Park
a2sys
Minsoo Rhu
Minsoo Rhu
KAIST
Computer ArchitectureVLSIMachine LearningComputer Vision