🤖 AI Summary
This study addresses the limited interpretability of expert routing mechanisms in Mixture-of-Experts (MoE) models and the shortcomings of the prevailing "single-domain specialization" assumption by proposing a novel "superimposed specialization" paradigm. To this end, we introduce RouterInterp, a method that integrates sparse autoencoders, feature prediction, and natural language generation to elucidate the intrinsic relationship between sparse features and routing decisions while automatically producing natural language explanations. This approach enables precise mapping from broad domains to fine-grained features. Evaluated on the gpt-oss-20b model, RouterInterp achieves an approximate 65% improvement in detection accuracy over conventional baselines, substantially enhancing the interpretability of MoE architectures.
📝 Abstract
Sparse Mixture of Experts (MoE) models scale more efficiently than dense models by routing tokens to modular expert networks that are only active for processing a fraction of tokens. A leading hypothesis for the performance of MoE models is that each expert specialises in a single, coherent domain. However, interpretability efforts that assume this hypothesis have generally been unsuccessful. We propose and present evidence for an alternative account that we call the Superposed Specialisation Hypothesis (SSH): experts specialise in a disjoint union of fine-grained features rather than one broad domain. Leveraging the SSH, we introduce RouterInterp, a method for interpreting expert routing that identifies Sparse Autoencoder features most predictive of routing decisions and produces unified natural language explanations. On gpt-oss-20b, RouterInterp explains expert routing with ${\sim}65\%$ higher detection accuracy than prior token statistics based methods. This work provides a scalable method for generating more accurate explanations of expert routing and increases our understanding of a previously uninterpretable component of foundation models.