Cross-Entropy Guided Routing in Mixture-of-Experts Large Language Models

πŸ“… 2026-09-29
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitation of existing Mixture-of-Experts (MoE) routers, which lack direct alignment with token-level error and thereby constrain the effectiveness of sparse routing optimization. To overcome this, we propose two error-aware routing mechanisms that incur no additional inference overhead. These methods achieve supervision by either predicting expert errors or aligning routing affinities with cross-entropy loss. By incorporating Itakura-Saito divergence, exponential negative log-likelihood, and Top-K strategies, the routing affinities are directly aligned with the model’s optimization objective. Evaluated on Granite baselines, our approach yields an average accuracy improvement of 2.3%, with gains reaching 2.94% on the ARC-Challenge benchmark, demonstrating efficient and precise sparse routing optimization.
πŸ“ Abstract
Sparse mixture-of-experts (MoE) large language models scale model capacity by routing each token to a small subset of experts. Their routers are regularized with load balancing terms and learn affinity scores through the language-model objective. However, these objectives do not provide direct alignment between routing affinities and token-level error. We introduce token-error supervision for sparse routing in two forms. The first form predicts an error score per expert. The affinity-weighted aggregate of these scores is aligned to the next-token cross-entropy loss, while the individual scores attenuate affinity before top-$K$ selection. The second directly aligns the router's affinities to the model's objective without requiring an additional head or inference-time modification. Both formulations use the Itakura--Saito divergence or an exponential negative log-likelihood for aligning affinities and token errors. Across two sparse MoE backbones and four multiple-choice question-answering benchmarks, we evaluate both supervision mechanisms. On Granite, our method improves accuracy by approximately 2.3 percentage points on average over a parameter-matched routing baseline. With stronger supervision, the gain on ARC-Challenge reaches 2.94 points. Both mechanisms preserve the native sparse execution budget and aggregation policy. Our code is available in the supplementary materials.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
Sparse Routing
Token-level Error
Cross-Entropy
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Experts
Cross-Entropy Guided Routing
Token-Error Supervision
Itakura-Saito Divergence
Sparse Routing