🤖 AI Summary
Gaussian processes (GPs) suffer from cubic time complexity $O(N^3)$ and quadratic memory cost $O(N^2)$, limiting scalability to large-scale or nonstationary data. To address this, we propose the Mixture of Gaussian Process Experts (MoE-GP) model, capable of capturing nonstationarity, heteroscedasticity, and discontinuities. We introduce the first nested sequential Monte Carlo (SMC²) inference framework for joint Bayesian inference over both the gating network and GP expert parameters. Our approach preserves full parallelizability while significantly improving posterior estimation accuracy and stability—reducing variance compared to standard importance sampling. Experiments demonstrate strong robustness and scalability on complex temporal and spatial datasets where conventional stationary GPs fail. MoE-GP establishes a novel paradigm for scalable, nonstationary GP modeling.
📝 Abstract
Gaussian processes are a key component of many flexible statistical and machine learning models. However, they exhibit cubic computational complexity and high memory constraints due to the need of inverting and storing a full covariance matrix. To circumvent this, mixtures of Gaussian process experts have been considered where data points are assigned to independent experts, reducing the complexity by allowing inference based on smaller, local covariance matrices. Moreover, mixtures of Gaussian process experts sub-stantially enrich the model’s flexibility, allowing for behaviors such as non-stationarity, heteroscedasticity, and discontinuities. In this work, we construct a novel inference approach based on nested sequential Monte Carlo samplers to simultaneously infer both the gating network and Gaussian process expert parameters. This greatly improves inference compared to importance sampling, particularly in settings when a stationary Gaussian process is inappropriate, while still being thoroughly parallelizable.