Conditional Capacity and Routing in Mixture-of-Experts Particle Transformers

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear trade-off between parameter capacity and computational cost in Mixture-of-Experts (MoE) architectures for particle physics Transformers. Based on the JetClass-II classification task, it employs dynamic routing algorithms and auxiliary loss optimization to systematically investigate how varying expert counts and routing strategies affect model performance, strictly decoupling memory capacity from active computation. The findings reveal that a Top-1 MoE configuration avoiding token dropping surpasses dense baselines without increasing nominal compute, whereas activating multiple experts improves accuracy at significantly higher computational overhead. Furthermore, the relationship between routing structure and performance is shown to be non-monotonic. This work provides essential theoretical foundations and practical guidelines for designing efficient MoE models in high-energy physics applications.
📝 Abstract
Mixture-of-Experts (MoE) models can increase parameter capacity without proportionally increasing active computation, but it is unclear how this trade-off behaves in particle-physics transformers. We study dense and MoE Particle Transformers on 188-class JetClass-II, varying expert count, routing capacity, top-K, and auxiliary loss. We find that, when token dropping is avoided, top-1 MoE models improve over the dense baseline at nearly unchanged nominal forward compute, while further increasing the number of stored experts produces little additional accuracy gain. Activating multiple experts per token yields additional predictive improvements at higher computational cost. Routing analyses show that expert assignments become more strongly associated with particle identity and kinematics in some configurations, but this structure does not increase monotonically with classification performance. These results highlight the need to distinguish stored parameter capacity, active computation, routing capacity, and routing organization when evaluating sparse expert models for jet classification. Code and experiment configurations are available at https://github.com/kpendiyala/MPT.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
Particle Transformer
Jet Classification
Conditional Capacity
Routing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Experts
Particle Transformer
Routing Capacity
Jet Classification
Sparse Models
K
Kaushik Pendiyala
University of California Davis, Davis, CA, USA
H
Haris Zia
University of California San Diego, San Diego, CA, USA
T
Trevin Lee
University of California San Diego, San Diego, CA, USA
T
Timothy Legge
University of California San Diego, San Diego, CA, USA
A
Alejandro J. De Leon
University of California San Diego, San Diego, CA, USA
Zihan Zhao
Zihan Zhao
Shanghai Jiao Tong University
NLP
A
Aaron Wang
University of Illinois at Chicago, Chicago, IL, USA
A
Abhijith Gandrakota
Fermi National Accelerator Laboratory, Batavia, IL, USA
Jennifer Ngadiuba
Jennifer Ngadiuba
Wilson Fellow, Fermilab
experimental high-energy physicsdata sciencedeep learningartificial intelligenceFPGAs
Richard Cavanaugh
Richard Cavanaugh
Professor of Physics and Senior Scientist, University of Illinois Chicago and Fermilab
Particle Physics
J
Javier Duarte
University of California San Diego, San Diego, CA, USA