🤖 AI Summary
This study addresses the unclear trade-off between parameter capacity and computational cost in Mixture-of-Experts (MoE) architectures for particle physics Transformers. Based on the JetClass-II classification task, it employs dynamic routing algorithms and auxiliary loss optimization to systematically investigate how varying expert counts and routing strategies affect model performance, strictly decoupling memory capacity from active computation. The findings reveal that a Top-1 MoE configuration avoiding token dropping surpasses dense baselines without increasing nominal compute, whereas activating multiple experts improves accuracy at significantly higher computational overhead. Furthermore, the relationship between routing structure and performance is shown to be non-monotonic. This work provides essential theoretical foundations and practical guidelines for designing efficient MoE models in high-energy physics applications.
📝 Abstract
Mixture-of-Experts (MoE) models can increase parameter capacity without proportionally increasing active computation, but it is unclear how this trade-off behaves in particle-physics transformers. We study dense and MoE Particle Transformers on 188-class JetClass-II, varying expert count, routing capacity, top-K, and auxiliary loss. We find that, when token dropping is avoided, top-1 MoE models improve over the dense baseline at nearly unchanged nominal forward compute, while further increasing the number of stored experts produces little additional accuracy gain. Activating multiple experts per token yields additional predictive improvements at higher computational cost. Routing analyses show that expert assignments become more strongly associated with particle identity and kinematics in some configurations, but this structure does not increase monotonically with classification performance. These results highlight the need to distinguish stored parameter capacity, active computation, routing capacity, and routing organization when evaluating sparse expert models for jet classification. Code and experiment configurations are available at https://github.com/kpendiyala/MPT.