MASKerade: Token-Routed Mask Experts for Dense-to-MoE Upcycling

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the severe storage and computational redundancy caused by independently duplicated expert weights when upgrading dense models to Mixture-of-Experts (MoE) architectures. To this end, it proposes MASKerade, which for the first time defines experts as learned subnetworks over a frozen pretrained FFN. Specifically, experts are instantiated via learned sparse binary masks and dynamically composed through token-level routing. By jointly optimizing the routing and mask parameters, the method supports various sparsity patterns, including 2:4 semi-structured sparsity, without introducing separate expert parameter matrices. Evaluated on Qwen and Gemma backbones, MASKerade achieves MoE-equivalent capacity at the arithmetic cost of a single dense forward pass. Furthermore, it outperforms baselines across five vision-language benchmarks, demonstrating both the efficiency and practicality of this mask-learning paradigm.
📝 Abstract
Sparsely activated Mixture-of-Experts (MoE) models increase model capacity without a proportional increase in per-token computation. Dense-to-MoE upcycling reuses pretrained dense models to construct such systems, commonly by copying feed-forward networks (FFNs) into independently trained experts. We introduce MASKerade, a dense-to-MoE training method that instead learns experts as sparse subnetworks of a frozen pretrained FFN. Each expert is defined by a learned binary mask, and a token-level router selects which masked FFNs to execute and combine. The router and mask scores are optimized jointly, while the underlying FFN weight values remain unchanged. This formulation supports neuron-structured, semi-structured, and unstructured experts within the same routing architecture. Our main configuration uses four 2:4 experts with top-2 routing, where two half-dense expert passes have the nominal FFN arithmetic of one dense pass, without requiring independent expert weight matrices. On five vision-language benchmarks with Qwen and Gemma backbones, this configuration achieves the highest performance among the compared baselines. Comparisons across mask granularities, routing interventions, and compute-matched controls distinguish the effects of learned connectivity from expert activation count. These results establish mask learning over frozen weights as a practical alternative for constructing token-routed MoE experts.
Problem

Research questions and friction points this paper is trying to address.

Dense-to-MoE upcycling
Mixture-of-Experts
sparse subnetworks
frozen pretrained weights
token routing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dense-to-MoE Upcycling
Mixture-of-Experts
Learned Binary Masks
Token-Routed Experts
Sparse Subnetworks
💼 Related Jobs
No related jobs found.
Mingyuan Zhang
Mingyuan Zhang
Northeastern University
Sparse NetworkLLM Post-trainingMulti-modal LearningMulti-view Learning
Yue Bai
Yue Bai
Northwestern University, Northeastern University
Multi-modal learningSparse network trainingMask learning
Zhongruo Wang
Zhongruo Wang
Amazon
Y
Yupin Huang
Department of Electrical and Computer Engineering, Northeastern University
Y
Yiyang Huang
Department of Electrical and Computer Engineering, Northeastern University
H
Hailing Wang
Department of Electrical and Computer Engineering, Northeastern University
Huimin Zeng
Huimin Zeng
Northeastern University
computer vision
Y
Yun Fu
Department of Electrical and Computer Engineering, Northeastern University; Khoury College of Computer Science, Northeastern University