🤖 AI Summary
This study investigates the impact of expert granularity on model performance within dense Mixture-of-Experts (MoE) architectures under a fixed computational budget. Theoretically, we demonstrate that dense MoE is equivalent to a standard feed-forward network (FFN) with token-dependent scaling, revealing a non-monotonic effect of fine-grained experts. Empirically, employing SwiGLU activations, Softmax gating, and fixed-width FFNs, we systematically compare varying numbers of activated experts in terms of validation loss and routing behavior. Our findings indicate that activating K=2 experts yields optimal performance, significantly outperforming the baseline, whereas higher granularity paradoxically degrades results. These observations confirm the critical role of soft gating mechanisms under specific architectural configurations.
📝 Abstract
Sparse Mixture-of-Experts (MoE) models combine learned routing with selection from a large expert pool. We isolate the contribution of dynamic expert combination using a dense analogue: $K$ SwiGLU experts, all active for every token and combined by a softmax gate, at fixed total FFN width. With no larger pool or discrete selection, the dense baseline is the $K=1$ case. Validation loss varies non-monotonically with $K$: $K=2$ improves over the baseline by $0.0048$, whereas $K=4$ and $K=6$ worsen it by $0.0053$ and $0.0197$, respectively. Routing generally remains soft and expert usage balanced, except in the first layer of the $K=4$ model, where routing is nearly one-hot. Forcing this gate to uniform increases loss by $2.4$ nats on a diagnostic subset, indicating that its concentrated routing is functionally important. We further show that the architecture is exactly a dense SwiGLU with token-dependent, simplex-constrained group scaling and contains the dense baseline in its function class. These single-run results suggest a granularity sweet spot for always-active, softly gated FFNs at fixed width, with $K=2$ performing best among the configurations tested.