SPEAR: A Spectral-Disentangled MoE Neural Operator with Knowledge-Guided Expert Aggregation for Large-Scale PDE Pretraining

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of knowledge interference and expert redundancy in Mixture-of-Experts (MoE) architectures for large-scale PDE pretraining by proposing the SPEAR model. Specifically, this work introduces a novel frequency-domain decoupling mechanism to disentangle feature frequencies, alongside a routing-preference-based expert similarity metric and a knowledge-guided aggregation strategy that effectively balances dynamic deduplication with specialized learning. Experimental evaluations across twelve PDE datasets demonstrate that SPEAR achieves state-of-the-art performance. Notably, it maintains or improves predictive accuracy even when the number of experts is halved, thereby successfully reconciling computational efficiency with strong generalization capabilities.
📝 Abstract
Large-scale pre-training has improved the generalization of neural operators across diverse PDEs. However, existing PDE foundation models still struggle with heterogeneous dynamics, where shared representations may cause knowledge interference, while mixture-of-experts (MoE) architectures suffer from increasing expert redundancy. We propose SPEAR, a spectral-disentangled MoE neural operator with knowledge-guided expert aggregation for large-scale PDE pre-training. SPEAR decouples latent features into low- and high-frequency components, enabling shared modeling of transferable dynamics and specialized learning of PDE-specific patterns. To address expert redundancy, we design a knowledge-guided expert aggregation strategy that measures expert similarity from dataset-specific learned knowledge and routing preferences, enabling the identification and consolidation of similar experts. Experiments on twelve PDE datasets and multiple downstream benchmarks demonstrate superior performance in pre-training, fine-tuning, and transfer learning. Furthermore, our aggregation strategy reduces the number of experts by 50\% while maintaining or improving prediction accuracy, achieving a balance between model efficiency and generalization for PDE foundation models.
Problem

Research questions and friction points this paper is trying to address.

PDE foundation models
heterogeneous dynamics
knowledge interference
expert redundancy
mixture-of-experts
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spectral-Disentangled MoE
Neural Operator
Knowledge-Guided Expert Aggregation
PDE Pretraining
Expert Redundancy
🔎 Similar Papers
2023-10-19International Conference on Machine LearningCitations: 5
Dengdi Sun
Dengdi Sun
Anhui University
Machine LearningComputer Vision
X
Xiaoya Zhou
School of Artificial Intelligence, Anhui University, Hefei 230601, China
X
Xiao Wang
School of Computer Science and Technology, Anhui University, Hefei 230601, China
W
Wanli Lyu
School of Computer Science and Technology, Anhui University, Hefei 230601, China
Jin Tang
Jin Tang
Anhui University
Computer visionintelligent video analysis
B
Bin Luo
School of Computer Science and Technology, Anhui University, Hefei 230601, China