Looping Beyond Twice: A Scalable Recipe for Looped Mixture-of-Experts

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of effectively scaling recurrence in Mixture-of-Experts (MoE) models, which are fundamentally constrained by the curse of depth and expert selection collapse. To overcome these limitations, this work proposes the LOOM framework, which introduces residual scaling, input reinjection, and layer-wise independent routers. By stabilizing recurrent states and diversifying expert routing, LOOM transcends the conventional bottleneck that restricts recurrence to merely two iterations. At a 1.7B parameter scale, the proposed method achieves stable nine-step recurrence, reducing perplexity to 7.77 and improving zero-shot accuracy to 47.7%. These results significantly outperform existing baselines, establishing a new paradigm for the efficient recurrent scaling of large-scale MoE models.
📝 Abstract
Looped Transformers introduce recurrent depth as a new scaling axis for LLMs: by repeatedly applying shared Transformer blocks, they increase effective depth without increasing parameter count. However, the benefits of looping remain unclear for large MoE LLMs under FLOPs-matched comparisons. The main reason is that the gains from additional iterations diminish quickly and can even turn into degradation, so the extra FLOPs spent on looping yield little substantial improvement. Consequently, prior work typically settles on two loops. We identify two main obstacles to scaling looped MoE. First, looping inherits and amplifies the curse of depth: hidden-state variance grows with each iteration as residual updates accumulate, which destabilizes deep recurrence and causes representations to drift. Second, looped MoE suffers from expert selection collapse: routers repeatedly select the same experts across loops, so extra iterations add computation without adding computational diversity. Guided by this diagnosis, we propose LOOM, built on a single principle: each loop should contribute new computation while keeping the recurrent state stable. LOOM stabilizes recurrence by scaling residual updates to bound variance growth and re-injecting the input embedding at every loop, and diversifies it through per-loop routers that engage different experts and a Looping Residual that carries earlier outputs forward. Experiments across 100M-1.7B models show stable scaling to 9-12 loops. Under near-iso-FLOP, the 700M model performs best at 5 loops, reducing perplexity from 18.36 to 16.54 and improving average zero-shot accuracy from 38.84% to 39.53% over the non-looped baseline. Without FLOP matching, the 1.7B model trained on 60B tokens peaks at 9 loops, reducing perplexity from 9.62 to 7.77 and improving average zero-shot accuracy from 42.4% to 47.7%. Code is available https://github.com/hed-ucas/LOOM.
Problem

Research questions and friction points this paper is trying to address.

Looped Transformers
Mixture-of-Experts
Curse of Depth
Expert Selection Collapse
Recurrent Stability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Looped Mixture-of-Experts
LOOM
recurrent depth scaling
expert selection collapse
hidden-state variance stabilization