🤖 AI Summary
This study addresses the limitation of existing scaling laws that model recurrence or sparsity in isolation, which hinders the principled design of recurrent Mixture-of-Experts (MoE) architectures. We propose the first scaling law that jointly characterizes recurrent depth and sparse capacity, formulating parameter gains through bounded mappings to unify dense and MoE models as special cases. Validated via trillion-token-scale training, our formulation accurately predicts model loss. Compared to non-recurrent MoE baselines with twice the size, the proposed approach achieves approximately 3× efficiency in active parameters and 2× efficiency in total parameters, while delivering superior performance on reasoning tasks. These findings establish a theoretical foundation for architecture design under resource constraints.
📝 Abstract
Looped transformers and Mixture-of-Experts (MoE) offer complementary routes to efficient scaling: recurrence increases computational depth at fixed parameters, while MoE sparsity expands total capacity at fixed active compute. Yet existing scaling laws model recurrence or sparsity in isolation. In this work, we introduce Loop Scaling Laws, the first scaling law to jointly model recurrence and sparsity alongside model size and data. At its core is a bounded, sparsity-conditional recurrence mapping that characterizes the effective-parameter gain from looping and how sparsity raises this gain. The laws predict the held-out loss of looped models more accurately than prior alternatives, and recover the standard dense and MoE scaling laws as special cases. Beyond prediction, the fitted laws provide a principled foundation for designing looped MoE models under compute and memory constraints. Downstream evaluations further demonstrate the complementary benefits of the two axes: sparsity delivers ~3x active-parameter efficiency, recurrence yields ~2x total-parameter efficiency on reasoning, and joint scaling further advances the performance frontier. As a practical extension, we show these gains hold at trillion-token scale: at matched training compute, a looped MoE with law-derived recurrence matches a ~2x larger non-looped MoE on the reasoning benchmarks, while enabling test-time scaling through recurrence.