🤖 AI Summary
This study addresses the challenges of knowledge interference and expert redundancy in Mixture-of-Experts (MoE) architectures for large-scale PDE pretraining by proposing the SPEAR model. Specifically, this work introduces a novel frequency-domain decoupling mechanism to disentangle feature frequencies, alongside a routing-preference-based expert similarity metric and a knowledge-guided aggregation strategy that effectively balances dynamic deduplication with specialized learning. Experimental evaluations across twelve PDE datasets demonstrate that SPEAR achieves state-of-the-art performance. Notably, it maintains or improves predictive accuracy even when the number of experts is halved, thereby successfully reconciling computational efficiency with strong generalization capabilities.
📝 Abstract
Large-scale pre-training has improved the generalization of neural operators across diverse PDEs. However, existing PDE foundation models still struggle with heterogeneous dynamics, where shared representations may cause knowledge interference, while mixture-of-experts (MoE) architectures suffer from increasing expert redundancy. We propose SPEAR, a spectral-disentangled MoE neural operator with knowledge-guided expert aggregation for large-scale PDE pre-training. SPEAR decouples latent features into low- and high-frequency components, enabling shared modeling of transferable dynamics and specialized learning of PDE-specific patterns. To address expert redundancy, we design a knowledge-guided expert aggregation strategy that measures expert similarity from dataset-specific learned knowledge and routing preferences, enabling the identification and consolidation of similar experts. Experiments on twelve PDE datasets and multiple downstream benchmarks demonstrate superior performance in pre-training, fine-tuning, and transfer learning. Furthermore, our aggregation strategy reduces the number of experts by 50\% while maintaining or improving prediction accuracy, achieving a balance between model efficiency and generalization for PDE foundation models.