X-MoD: Practical Scaling Laws for Sparse-Depth Routing Beyond Mixture-of-Depths

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of sparse depth scaling and capacity coupling in Mixture-of-Depths (MoD) architectures by proposing X-MoD. The method decouples token sparsity from anchor stride, enabling total parameter growth while maintaining constant activation capacity. It formulates sparse depth routing as a conditional architecture design problem for the first time, establishing an interpretable scaling law to decompose performance gains. Furthermore, dense anchors, variance-scaled gating, and depth-wise token balancing techniques are introduced to optimize training. Pretraining experiments demonstrate that this framework accurately predicts validation loss across diverse configurations and significantly outperforms both Dense and MoE baselines.
📝 Abstract
Mixture-of-Depths (MoD) enables conditional computation across Transformer depth by routing only a subset of tokens through selected layers, but its original one-sparse--one-dense alternation tightly couples total capacity to active capacity and limits sparse-depth scaling. We introduce X-MoD, a scalable sparse-depth architecture that decouples token sparsity from anchor stride, allowing total parameter count to grow while keeping active-equivalent capacity nearly fixed. To make deep sparse routing trainable, X-MoD combines dense anchors with variance-scaled layer-wise gating and depth-wise token balancing. To make this regime analyzable and usable, we formulate sparse-depth routing as a conditional architecture-design problem: given compute, context length, and active-equivalent backbone size, how should the routing configuration be chosen? We develop a practical scaling-law framework by fitting X-MoD relative to FLOP-matched dense baselines, yielding an interpretable law that decomposes performance into sparse-capacity gain, sparse-context correction, and anchor-stride interaction. The law predicts validation loss across routing configurations and reveals how context length, model scale, and anchor stride shape sparse-depth performance. We validate the architecture and law through pretraining sweeps, held-out scaling-law prediction, ablations, downstream evaluations, and comparisons with Dense, MoD, and representative MoE baselines.
Problem

Research questions and friction points this paper is trying to address.

Sparse-Depth Routing
Mixture-of-Depths
Scaling Laws
Conditional Computation
Token Sparsity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sparse-Depth Routing
Mixture-of-Depths
Scaling Laws
Conditional Computation
Token Sparsity
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Bowen Dong
Bowen Dong
Tsinghua University
Y
Yilong Fan
Tianjin University
T
Tengyu Pan
Tsinghua University
Y
Yike Zhang
Tsinghua University
Zhenyu Li
Zhenyu Li
Tsinghua University
Language ModelReasoningLong ContextReinforcement Learning
Z
Zijian Zhang
Tianjin University
X
Xuewei Li
Tianjin University
M
Mei Yu
Tianjin University
Jianyong Wang
Jianyong Wang
Tsinghua University
Data MiningKnowledge GraphMedical Data Mining