🤖 AI Summary
This study addresses the semantic routing fragmentation in traditional Mixture-of-Experts (MoE) for video diffusion models caused by uniformity constraints. We propose SplitMoE, a novel architecture that introduces a role-disentangled sparsity mechanism to explicitly separate semantic and general experts, thereby breaking the uniformity trap. Furthermore, we incorporate prototype-guided routing with push-pull regularization to effectively overcome routing imbalances induced by spatiotemporal redundancy in visual data, achieving precise semantic clustering. Experimental results demonstrate that, under an equivalent parameter budget, our method significantly improves convergence speed, routing coherence, and video generation quality. This work provides a new pathway toward modality-aware efficient scaling for video diffusion models.
📝 Abstract
Mixture-of-Experts (MoE), popularized by large language models, is a promising paradigm for scaling visual generative models. However, conventional token-wise MoE routes tokens independently within a homogeneous expert pool and regularizes expert usage toward uniformity, making it poorly matched to video data that is spatiotemporally redundant and semantically long-tailed. We show that existing visual MoEs fall into a uniformity trap: semantically under-organized routing, compounded by uniform expert-usage regularization, scatters coherent patches across disparate experts, causing routing fragmentation and structural distortion. To address this, we propose SplitMoE, a split-role sparse architecture that breaks the shackles of uniformity. To accommodate the inherent semantic imbalance, we explicitly bifurcate the expert pool into semantic experts and generic experts, with semantic experts capturing high-level semantic abstraction and generic experts preserving residual visual information and flexible generative capacity. Leveraging prototype-guided routing and pull-push regularization, SplitMoE enables tokens to cluster naturally by semantic attributes rather than arbitrary balancing constraints. Extensive results show that under an equivalent activated-parameter budget, SplitMoE outperforms traditional load-balanced MoEs in convergence speed, routing coherence, and video generation quality across standard benchmarks. By revealing an emergent coarse-to-fine denoising logic, SplitMoE provides the community with a modality-aware scaling path, serving as a critical reference for building large-scale video world models.