🤖 AI Summary
This study addresses the unclear performance gap between dense and sparse Mixture-of-Experts (MoE) models at extremely small scales (<25M parameters). Under a unified LLaMA-style pretraining setup, the authors systematically compare both architectures while fixing all hyperparameters except model width, aligning either active or total parameter counts. To stabilize sparse training, they incorporate Switch-style load balancing and router z-loss. Results show that when active parameters are matched, MoE achieves lower validation loss than its dense counterpart (1.5788 vs. 1.6545). However, under total parameter alignment—reflecting equal storage footprint—the dense model slightly outperforms MoE (1.5608 vs. 1.5788), revealing that MoE struggles to surpass dense models of equivalent capacity at very small scales.
📝 Abstract
We study dense and mixture-of-experts (MoE) transformers in a tiny-scale pretraining regime under a shared LLaMA-style decoder training recipe. The sparse model replaces dense feed-forward blocks with Mixtral-style routed experts. Dense baselines are modestly width-resized to tightly match either active or total parameter budgets, while tokenizer, data, optimizer, schedule, depth, context length, normalization style, and evaluation protocol are held fixed. Our best sparse recipe uses four experts, top-2 routing, Switch-style load balancing, and router z-loss. In a three-seed full-data comparison, the dense active-match model reaches 1.6545 +/- 0.0012 best validation loss, the MoE reaches 1.5788 +/- 0.0020, and the dense total-match model reaches 1.5608 +/- 0.0025. This yields a matched-active gap of 0.0758 +/- 0.0021 in the MoE's favor and a matched-total gap of 0.0180 +/- 0.0020 in the dense model's favor. Across training, the matched-active advantage grows while the matched-total dense advantage narrows sharply. In this sub-25M-parameter regime, MoE therefore improves validation loss under active-parameter matching but does not surpass dense training at equal total stored capacity.