🤖 AI Summary
This study addresses the limitation of existing Mixture-of-Experts (MoE) routing mechanisms that fail to leverage historical distributions from preceding layers, thereby constraining training efficiency. To overcome this, we propose HERO-MoE, a framework that reuses prior-layer routing distributions as informative priors. Specifically, it optimizes current routing decisions through auxiliary-loss-free historical routing residual injection and adaptive scale-preserving fusion. This design remains fully compatible with standard top-k sparse dispatching while requiring minimal architectural modifications. Experimental results on an 8B-parameter model demonstrate that HERO-MoE reduces the final loss to 1.6184, incurring only marginal overheads of 0.44% in GPU memory and 0.64% in FLOPs. Consequently, the proposed approach significantly accelerates training convergence at a negligible computational cost.
📝 Abstract
Mixture-of-Experts (MoE) architectures have become a standard way to scale model capacity while keeping computation sparse, yet routing remains a key determinant of MoE quality and training behavior. Prior empirical studies suggest that MoE routing reflects input semantics and upstream computation across depth, but standard routers do not explicitly use the routing distributions produced by preceding layers. We propose HERO-MoE, Historical Expert ROuting with Scale-Preserving Fusion, a routing framework that injects historical routing priors into MoE routers by reusing detached, dense routing distributions collected from preceding MoE layers. The key idea is simple: HERO-MoE preserves the original token-conditioned routing branch and adds a residual historical routing contribution before the standard softmax and top-$k$ dispatch. To stabilize this historical signal, HERO-MoE introduces a scale-preserving fusion mechanism that matches the magnitude of historical routing memory to the current hidden representation and accounts for the number of visible historical layers, without introducing an auxiliary routing loss or a fusion-specific tuning parameter. By reusing routing distributions already computed by preceding MoE layers, HERO-MoE improves training-loss reduction with modest end-to-end overhead. The resulting router remains compatible with standard sparse dispatch, including top-$k$ and group-limited routing, and can be inserted into existing MoE backbones with minimal architectural changes. Experiments on an approximately 8B-parameter MoE model with 0.5B active parameters, trained from scratch on 100B tokens, show that HERO-MoE reduces the final loss from 1.6393 to 1.6184, while peak memory and FLOPs increase by only 0.44\% and 0.64\%, respectively.