dual-stream moe

Designs and implements neural architectures that split processing into two interleaved or parallel token streams governed by a sparsely-activated mixture-of-experts, building gating and routing mechanisms to allocate capacity and route tokens into capability-specific expert pathways; evaluates and analyzes how these dual streams share joint context, decouple understanding and generation flows, and affect training dynamics, routing stability, and inference efficiency.

dual-streammoe

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.41
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Neural Inhibition Improves Dynamic Routing and Mixture of Experts

Jul 03, 2025
WY
Will Y. Zou
🏛️ Angle.ac | University of Toronto

Dynamic routing in Mixture-of-Experts (MoE) models is vulnerable to redundant neuron activations, leading to biased expert selection and insufficient expert diversity. Method: We propose a neural suppression–enhanced dynamic routing mechanism that applies learnable suppression signals to redundant neuron populations within the shared feature space prior to routing decisions, explicitly attenuating collinear responses to improve discriminability and specialization of expert path selection. Contribution/Results: Unlike prior MoE approaches, this work is the first to systematically demonstrate the routing-quality benefits of neural suppression and integrate it end-to-end into Transformer-like architectures without increasing parameter count. Experiments across multiple NLP and vision benchmarks show consistent improvements: +1.2–2.8% accuracy gains, −37% reduction in routing variance, and enhanced expert utilization balance and task adaptability—establishing a novel paradigm for efficient, diverse sparse modeling.

Boosting performance in Mixture-of-Experts and transformer modelsEnhancing specialized expert path selection via inhibitionImproving dynamic routing models with neural inhibition

This work systematically dissects the multidimensional design space of Mixture-of-Experts (MoE) architectures in large language models, moving beyond conventional generational narratives. It introduces a five-dimensional analytical framework encompassing expert granularity, topology, routing flexibility, load-balancing scope, and execution structure, and constructs a dependency graph to elucidate the coupling mechanisms across four control planes: expert topology, routing, load balancing, and expert parallelism. The framework’s validity is empirically demonstrated through iso-budget pretraining experiments integrating algorithmic innovations—such as Top-k routing, shared and fine-grained experts, and dynamic expert composition—with system-level optimizations including token dispatch, device placement, and all-to-all communication. The study further distills key open challenges for the future development of MoE systems.

expert parallelismexpert topologyload balancing

This work addresses the routing collapse and expert deadlocks that commonly afflict Token-Choice sparse Mixture-of-Experts (MoE) architectures in video diffusion Transformers, which severely limit expert diversity utilization. Starting from a 5-billion-parameter dense model, the authors formulate three principles for converting dense networks to MoE. Through temporal routing analysis of 65 million tokens, they reveal that deadlocked layers follow a U-shaped distribution across the network depth and propose a “functional redundancy” hypothesis to explain this phenomenon. Building on these insights, they integrate expert cloning, zero-initialized gating, auxiliary losses, and enhanced router designs—including linear, MLP, and cross-attention variants—to effectively mitigate bfloat16 precision pitfalls. Their approach alleviates single-expert deadlocks in approximately two-thirds of network layers, endows the model with partial self-recovery capability, delineates the capacity limits of the Token-Choice paradigm, and outlines a three-stage roadmap toward unified vision models and ultimately world models.

Diffusion TransformersRouting CollapseSelective Deadlock

Load Balancing Mixture of Experts with Similarity Preserving Routers

Jun 16, 2025
NO
Nabil Omi
🏛️ University of Washington | Microsoft Research | Allen Institute for AI

Sparse Mixture-of-Experts (MoE) models often suffer from capacity waste and performance degradation due to routing bias toward a small subset of experts. Conventional load-balancing methods enforce uniform expert utilization but risk undermining semantic coherence, leading to knowledge redundancy across experts. To address this, we propose a similarity-preserving load-balancing mechanism: a differentiable routing loss grounded in token embedding similarity, which encourages semantically similar tokens to be consistently routed to the same expert—thereby jointly optimizing load distribution and routing consistency. Our approach requires no additional experts or auxiliary modules and integrates seamlessly into standard MoE training pipelines. Experiments demonstrate a 36% acceleration in convergence on benchmark tasks, substantial reduction in inter-expert knowledge redundancy, and improved model generalization and inference efficiency.

Ensures consistent expert assignment for similar inputsPrevents expert underutilization in sparse MoE modelsReduces redundant knowledge learning in expert routing

This work addresses the disconnect in existing Mixture-of-Experts (MoE) models between shared computation and dynamic routing, which overlooks the interdependence between reusable computation and residual expert requirements. The paper proposes UniF-MoE, a unified framework introducing a novel “shared-first, routed-later” mechanism: it first processes common features through a shared general-purpose module and then dynamically activates residual experts based on a shared-demand score and complementarity. Key innovations include key prototype selection, cumulative routing quality allocation, and Gram regularization to enhance routing sparsity and diversity, revealing a negative correlation between shared coverage and residual demand. Experiments demonstrate that UniF-MoE outperforms both static and dynamic MoE approaches on DomainBed and GLUE benchmarks while significantly reducing activated computation, inference latency, and memory footprint.

dynamic routingexpert capacityMixture-of-Experts

Latest Papers

What's happening recently
View more

This study addresses the substantial All-to-All communication overhead in Mixture-of-Experts (MoE) expert parallelism, which accounts for 45%–60% of training step time. By observing significant intra-layer and inter-layer correlations in expert selection during early pre-training stages, this work proposes a routing-correlation-based communication optimization method. Specifically, correlation-aware expert placement strategies and token shuffling mechanisms are designed within the Megatron-LM framework to effectively reduce cross-GPU data transfers. Experimental results demonstrate that the proposed approach reduces All-to-All communication latency by 1.16× to 2.63× and achieves up to a 1.41× end-to-end training step speedup, substantially improving the distributed training efficiency of MoE models.

All-to-All Communication OverheadDistributed TrainingExpert Parallelism

This work addresses the pervasive issue of deep routing collapse in large Mixture-of-Experts (MoE) models for low-resource languages, which leads to imbalanced expert utilization and constrained multilingual capabilities. The study reveals, for the first time, that this phenomenon stems from insufficient pretraining data rather than inherent linguistic properties. To diagnose multilingual capacity, the authors propose routing entropy and expert specialization as key indicators. Through balanced bilingual continual pretraining (CPT) and supervised fine-tuning (SFT) on Hebrew, Japanese, and other languages using both pure Transformer and Mamba-Transformer hybrid architectures, they demonstrate that CPT substantially increases routing entropy, encourages language-agnostic expert sharing, and consistently enhances downstream performance, whereas SFT yields limited gains. These findings underscore the critical role of data balance in scaling MoE models multilingually.

expert entropylow-resource languagesMixture-of-Experts

This work systematically investigates the interplay among key design dimensions in Mixture-of-Experts (MoE) architectures—such as the number of experts, expert granularity, heterogeneity, shared experts, and load balancing—through over 2,000 large-scale pretraining experiments. The study reveals that the number of experts and their granularity are the dominant factors governing model performance, while other design choices exert comparatively limited influence. Notably, increasing the total MoE parameters consistently enhances performance across all active parameter budgets, and the optimal expert size is determined solely by the number of active parameters. Furthermore, the effectiveness of dropless routing is empirically validated, demonstrating consistent performance gains.

expert countexpert granularityload balancing

Hot Scholars

DH

Dongchen Han

Tsinghua University
Computer VisionDeep Learning
YQ

Yanyuan Qiao

Postdoctoral Research Fellow, EPFL
Embodied-AIVision and LanguageMulti-modal Learning
XH

Xiaoshuai Hao

Beijing Academy of Artificial Intelligence,BAAI
vision and language
QL

Qixiu Li

Tsinghua University
Embodied AIComputer VisionMachine Learning
QZ

Qi Zhu

Nanjing University of Aeronautics and Astronautics
brain networkmedical image processingartificial intelligence