π€ AI Summary
This study addresses the limitation of existing Mixture-of-Experts (MoE) routers, which lack direct alignment with token-level error and thereby constrain the effectiveness of sparse routing optimization. To overcome this, we propose two error-aware routing mechanisms that incur no additional inference overhead. These methods achieve supervision by either predicting expert errors or aligning routing affinities with cross-entropy loss. By incorporating Itakura-Saito divergence, exponential negative log-likelihood, and Top-K strategies, the routing affinities are directly aligned with the modelβs optimization objective. Evaluated on Granite baselines, our approach yields an average accuracy improvement of 2.3%, with gains reaching 2.94% on the ARC-Challenge benchmark, demonstrating efficient and precise sparse routing optimization.
π Abstract
Sparse mixture-of-experts (MoE) large language models scale model capacity by routing each token to a small subset of experts. Their routers are regularized with load balancing terms and learn affinity scores through the language-model objective. However, these objectives do not provide direct alignment between routing affinities and token-level error. We introduce token-error supervision for sparse routing in two forms. The first form predicts an error score per expert. The affinity-weighted aggregate of these scores is aligned to the next-token cross-entropy loss, while the individual scores attenuate affinity before top-$K$ selection. The second directly aligns the router's affinities to the model's objective without requiring an additional head or inference-time modification. Both formulations use the Itakura--Saito divergence or an exponential negative log-likelihood for aligning affinities and token errors. Across two sparse MoE backbones and four multiple-choice question-answering benchmarks, we evaluate both supervision mechanisms. On Granite, our method improves accuracy by approximately 2.3 percentage points on average over a parameter-matched routing baseline. With stronger supervision, the gain on ARC-Challenge reaches 2.94 points. Both mechanisms preserve the native sparse execution budget and aggregation policy. Our code is available in the supplementary materials.