Dense Mixture-of-Experts as a Reparameterized Wide FFN: A Granularity Sweep at Fixed Compute

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the impact of expert granularity on model performance within dense Mixture-of-Experts (MoE) architectures under a fixed computational budget. Theoretically, we demonstrate that dense MoE is equivalent to a standard feed-forward network (FFN) with token-dependent scaling, revealing a non-monotonic effect of fine-grained experts. Empirically, employing SwiGLU activations, Softmax gating, and fixed-width FFNs, we systematically compare varying numbers of activated experts in terms of validation loss and routing behavior. Our findings indicate that activating K=2 experts yields optimal performance, significantly outperforming the baseline, whereas higher granularity paradoxically degrades results. These observations confirm the critical role of soft gating mechanisms under specific architectural configurations.
📝 Abstract
Sparse Mixture-of-Experts (MoE) models combine learned routing with selection from a large expert pool. We isolate the contribution of dynamic expert combination using a dense analogue: $K$ SwiGLU experts, all active for every token and combined by a softmax gate, at fixed total FFN width. With no larger pool or discrete selection, the dense baseline is the $K=1$ case. Validation loss varies non-monotonically with $K$: $K=2$ improves over the baseline by $0.0048$, whereas $K=4$ and $K=6$ worsen it by $0.0053$ and $0.0197$, respectively. Routing generally remains soft and expert usage balanced, except in the first layer of the $K=4$ model, where routing is nearly one-hot. Forcing this gate to uniform increases loss by $2.4$ nats on a diagnostic subset, indicating that its concentrated routing is functionally important. We further show that the architecture is exactly a dense SwiGLU with token-dependent, simplex-constrained group scaling and contains the dense baseline in its function class. These single-run results suggest a granularity sweet spot for always-active, softly gated FFNs at fixed width, with $K=2$ performing best among the configurations tested.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
Dense MoE
Feed-Forward Network
Granularity
Fixed Compute
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dense Mixture-of-Experts
Reparameterized FFN
SwiGLU
Soft Gating
Granularity Sweep
💼 Related Jobs
No related jobs found.
V
Vu Quang Hoang
University of Information Technology, Vietnam National University, Ho Chi Minh City, Viet Nam
N
Nghia Hieu Nguyen
University of Information Technology, Vietnam National University, Ho Chi Minh City, Viet Nam