HERO-MoE: Historical Expert Routing with Scale-Preserving Fusion

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing Mixture-of-Experts (MoE) routing mechanisms that fail to leverage historical distributions from preceding layers, thereby constraining training efficiency. To overcome this, we propose HERO-MoE, a framework that reuses prior-layer routing distributions as informative priors. Specifically, it optimizes current routing decisions through auxiliary-loss-free historical routing residual injection and adaptive scale-preserving fusion. This design remains fully compatible with standard top-k sparse dispatching while requiring minimal architectural modifications. Experimental results on an 8B-parameter model demonstrate that HERO-MoE reduces the final loss to 1.6184, incurring only marginal overheads of 0.44% in GPU memory and 0.64% in FLOPs. Consequently, the proposed approach significantly accelerates training convergence at a negligible computational cost.
📝 Abstract
Mixture-of-Experts (MoE) architectures have become a standard way to scale model capacity while keeping computation sparse, yet routing remains a key determinant of MoE quality and training behavior. Prior empirical studies suggest that MoE routing reflects input semantics and upstream computation across depth, but standard routers do not explicitly use the routing distributions produced by preceding layers. We propose HERO-MoE, Historical Expert ROuting with Scale-Preserving Fusion, a routing framework that injects historical routing priors into MoE routers by reusing detached, dense routing distributions collected from preceding MoE layers. The key idea is simple: HERO-MoE preserves the original token-conditioned routing branch and adds a residual historical routing contribution before the standard softmax and top-$k$ dispatch. To stabilize this historical signal, HERO-MoE introduces a scale-preserving fusion mechanism that matches the magnitude of historical routing memory to the current hidden representation and accounts for the number of visible historical layers, without introducing an auxiliary routing loss or a fusion-specific tuning parameter. By reusing routing distributions already computed by preceding MoE layers, HERO-MoE improves training-loss reduction with modest end-to-end overhead. The resulting router remains compatible with standard sparse dispatch, including top-$k$ and group-limited routing, and can be inserted into existing MoE backbones with minimal architectural changes. Experiments on an approximately 8B-parameter MoE model with 0.5B active parameters, trained from scratch on 100B tokens, show that HERO-MoE reduces the final loss from 1.6393 to 1.6184, while peak memory and FLOPs increase by only 0.44\% and 0.64\%, respectively.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
Routing mechanism
Historical routing priors
Model scaling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Experts
Historical Expert Routing
Scale-Preserving Fusion
Sparse Dispatch
Routing Prior
💼 Related Jobs
No related jobs found.
J
Junxiang Qiu
University of Science and Technology of China
Z
Zhengsu Chen
Huawei Inc.
Xinting Hu
Xinting Hu
Max Planck Institute for Informatics
Multimodal ReasoningContinual LearningSemi-Supervised Learning
Shuo Wang
Shuo Wang
University of Science and Technology of China
Computer VisionMultimedia
H
Hengheng Zhang
Huawei Inc.
Shaofeng Zhang
Shaofeng Zhang
Southern University of Science and Technology
Learn to Optimize
C
Changcheng Li
University of Science and Technology of China
B
Boyu Shi
Southeast University
Q
Qi Tian
Huawei Inc.