MoRE: Scaling mixture of experts with hardware-aware low-rank routing

πŸ“… 2026-09-28
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the computational bottleneck of linear routers in large-scale Mixture-of-Experts (MoE) models by proposing the MoRE architecture, which leverages low-rank matrix decomposition to substantially reduce routing overhead and support larger expert counts. Theoretically, we prove that a logarithmic rank suffices to preserve routing expressivity and load balancing, thereby overcoming conventional full-rank constraints. From an engineering perspective, efficient deployment is achieved through Triton fused kernels and hardware-aware optimizations. Experimental results demonstrate that this approach maintains inference performance while effectively enhancing model memorization capacity and knowledge-based question answering capabilities, ultimately enabling significant scaling of expert numbers.
πŸ“ Abstract
Mixture-of-Experts (MoE) layers are central to frontier language models, and recent architectures push toward more and smaller experts. In this regime, the standard linear router becomes a bottleneck: with $M$ experts and hidden dimension $h$, its per-token cost $Θ(Mh)$ dominates the MoE layer once $M$ is large. We introduce MoRE (Mixture of Rank-reduced-routed Experts), which factorizes the router weight matrix at rank $r$ and reduces the routing cost to $O((h + M)r)$. We prove that rank logarithmic in $M$ suffices for routing expressivity when the number of active experts is fixed, and is necessary up to precision factors. We also prove that logarithmic rank preserves load balance in a Gaussian memorization model, and training on a synthetic phonebook task shows that low rank does not hurt memorization. At matched active FLOPs, the factorization allows a factor of $Θ(h/r)$ more experts. To realize this gain in wall-clock time, we design a fused Triton kernel at inference that avoids expensive memory operations on HBM. Empirically, MoRE improves memorization on the phonebook task and performance on knowledge-intensive Q\&A benchmarks after pretraining, while matching reasoning ability. Code available at https://github.com/Matheart/MoRE_code.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
routing bottleneck
low-rank routing
hardware-aware inference
scalability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Experts
Low-rank routing
Hardware-aware optimization
Triton kernel
Scalability
πŸ”Ž Similar Papers
No similar papers found.