ACE: Adaptive Calibration-Free Expert Skipping for MoE-based LLMs

📅 2026-09-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出ACE框架,通过无训练、无校准的方法自适应跳过MoE架构中的低贡献专家,减少冗余计算,提高模型效率与性能。
📝 Abstract
Mixture-of-Experts (MoE) architectures provide an efficient paradigm for scaling large language models (LLMs), yet fixed top-k routing activates the same number of expert slots for every token, causing substantial redundant computation. Existing expert-skipping methods often rely on router confidence, calibration data, or additional training, and therefore cannot reliably estimate the actual contribution of routed experts. To this end, we propose ACE, a training-free, calibration-free, and checkpoint-preserving framework for token-adaptive expert skipping in MoE-based LLMs. ACE contains two complementary components: 1) Global Spectral Proxy (GSP), which estimates global transformation capacity from the coupled gate, up, and down projections together with RMSNorm scaling; and 2) Router-Conditioned Refinement (RCR), which constructs expert-specific direction prototypes from centered router weights and evaluates expert responses along routing-preferred directions. During inference, ACE combines both estimates with runtime router gates and skips an expert slot only when both views identify it as low-contribution, while always retaining the top-1 expert. All expert statistics are computed offline, leaving only table lookups and lightweight scalar operations online. Extensive experiments across three MoE-based LLMs and eight benchmarks demonstrate that ACE consistently outperforms existing static and dynamic baselines, with increasingly pronounced advantages under aggressive expert skipping. For instance, at a 50% skipping ratio on Qwen3.6-35B-A3B, ACE reduces WikiText-2 perplexity by 7.96% and improves average downstream accuracy by 4.15 percentage points over the strongest competing method.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
redundant computation
expert-skipping
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive Expert Skipping
Mixture-of-Experts
Global Spectral Proxy
Router-Conditioned Refinement
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zukang Xu
Z
Zhixiong Zhao
X
Xing Hu
Jiangyong Yu
Jiangyong Yu
houmo.ai
H
Houji Wen
J
Jun Li
Z
Zhe Jiang
D
Dawei Yang