Tight Clusters Make Specialized Experts

📅 2025-02-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
In high-dimensional sparse Mixture-of-Experts (MoE) models, conventional routers struggle to discern latent token clustering structures, resulting in slow convergence, poor robustness against data corruption, and degraded representation learning. To address this, we propose the Adaptive Clustering (AC) router: it employs a learnable feature-weighting mapping to project tokens into a latent space conducive to expert separation; and introduces, for the first time, an expert-level compactness-aware dynamic feature weighting mechanism—enabling each expert to specialize within semantically coherent subspaces. The method integrates clustering optimization theory, adaptive feature scaling, and token-level routing reparameterization. Evaluated on language modeling and image recognition tasks, the AC router achieves significant improvements in convergence speed, robustness, and overall performance, effectively mitigating routing failure induced by clustering unidentifiability in high-dimensional spaces.

Technology Category

Machine Learning: Mixture of Experts (MoE)Planning, Routing, and Scheduling: Planning with Language ModelsSearch and Optimization: Learning to Search

Application Category

User Modeling, Personalization and Recommendation: Federated recommendation systems and personalizationGraph Algorithms and Modeling for the Web: Graph embeddings and representation learning for Web-related graphsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
📝 Abstract
Sparse Mixture-of-Experts (MoE) architectures have emerged as a promising approach to decoupling model capacity from computational cost. At the core of the MoE model is the router, which learns the underlying clustering structure of the input distribution in order to send input tokens to appropriate experts. However, latent clusters may be unidentifiable in high dimension, which causes slow convergence, susceptibility to data contamination, and overall degraded representations as the router is unable to perform appropriate token-expert matching. We examine the router through the lens of clustering optimization and derive optimal feature weights that maximally identify the latent clusters. We use these weights to compute the token-expert routing assignments in an adaptively transformed space that promotes well-separated clusters, which helps identify the best-matched expert for each token. In particular, for each expert cluster, we compute a set of weights that scales features according to whether that expert clusters tightly along that feature. We term this novel router the Adaptive Clustering (AC) router. Our AC router enables the MoE model to obtain three connected benefits: 1) faster convergence, 2) better robustness to data corruption, and 3) overall performance improvement, as experts are specialized in semantically distinct regions of the input space. We empirically demonstrate the advantages of our AC router over baseline routing methods when applied on a variety of MoE backbones for language modeling and image recognition tasks in both clean and corrupted settings.
Problem

Research questions and friction points this paper is trying to address.

Optimizes token-expert routing in MoE models
Enhances clustering for better expert specialization
Improves robustness and performance in diverse tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive Clustering router
Optimal feature weights
Semantically specialized experts
💼 Related Jobs
No related jobs found.
S
Stefan K. Nielsen
FPT Software AI Center
R
R. Teo
Department of Mathematics, National University of Singapore
L
Laziz U. Abdullaev
Department of Mathematics, National University of Singapore
T
Tan M. Nguyen
Department of Mathematics, National University of Singapore