How Sparse Probability Maps Shape Mixture-of-Experts Routing

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether sparse probability mappings can achieve adaptive expert participation in Mixture-of-Experts (MoE) routing. Using 300M and 1B parameter MoE language models, we systematically compare Softmax with sparse routing mechanisms including Sparsemax, Entmax, Normmax, and Top-K. Our analysis reveals a co-adaptation phenomenon between probability mappings and router scores: routers collaboratively adjust score distributions to counteract the sparsity constraints imposed by the mapping. This indicates that the mapping alone cannot determine expert participation rates, necessitating the joint design of both the mapping and scoring mechanisms. Experiments demonstrate no significant differences in validation loss across mappings; however, sparse mappings substantially enhance robustness to variations in the number of active experts during inference.
📝 Abstract
Mixture-of-experts (MoE) routers typically apply softmax to the router scores and keep the top-K experts, making every token use exactly K experts. Sparsity-inducing probability maps such as sparsemax, alpha-entmax and normmax can adaptively assign exact zeros to selected experts, and therefore appear to offer token-dependent expert participation, even when using the same top-K machinery. In this work, we study whether and how this sparsity survives training. We train matched 300M and 1B top-2 MoE language models with softmax, 1.5-entmax, sparsemax and 2-normmax, and find that the maps behave very differently once trained: at 1B, entmax discards 30% less probability mass than softmax while almost never dropping a selected expert, sparsemax retains the most mass, and normmax routes 21% of tokens to a single expert. These outcomes are not properties of the maps alone. Each map drops a selected expert only when the gap between the two largest scores reaches a fixed threshold, and the trained routers differ in the score distribution they learn: the entmax router learns scores with roughly half the spread of softmax's, which keeps its top-2 gaps below its threshold, while sparsemax and normmax, which share the same threshold, learn different gap distributions and hence different participation. Routers thus co-adapt their scores to the map, and a map's capacity to produce zeros does not by itself determine expert participation. While none of the sparse maps improves validation loss over softmax, they make the trained models far less sensitive to selecting more experts at inference: sparsemax trained with K=2 loses 0.02 nats when run with K=8, where softmax loses 0.58. Our results indicate that adaptive MoE routing has to be designed around the joint behavior of the probability map and the learned scores, rather than around the map alone.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
Sparse Probability Maps
Routing
Sparsity
Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Experts
Sparse Probability Maps
Adaptive Routing
Co-adaptation
Inference Robustness
🔎 Similar Papers
No similar papers found.
T
Tomás Brogueira
Técnico, Universidade de Lisboa; INESC-ID; Instituto de Telecomunicações; ELLIS Unit Lisbon
M
Marcos Treviso
Técnico, Universidade de Lisboa; Instituto de Telecomunicações; ELLIS Unit Lisbon; Gandara AI
Miguel Couceiro
Miguel Couceiro
Full Professor at IST, U.Lisboa, INESC-ID
Knowledge discoveryAnalogy based reasoningDecision makingFair and explainable models