Score
Designs, implements, and evaluates approximate Gaussian process priors and inference methods that reduce compute and memory costs while preserving uncertainty quantification; builds sparse and hierarchical representations (e.g., inducing-point/sparse GP approximations, sparse precision matrices, hierarchical prior approximations) and scalable algorithms to enable GP inference on long sequences or large datasets.
Gaussian processes (GPs) suffer from cubic time and quadratic space complexity, hindering scalability to large datasets and modern parallel hardware. To address this, we propose a novel scalable GP inference framework that unifies iterative linear solvers with pathwise conditioning—replacing Cholesky-based inference with matrix-multiplication-dominated iterative linear system solving. This reformulation reduces memory complexity from $O(n^2)$ to $O(n)$, enables efficient GPU/TPU acceleration, and supports training and prediction on datasets with over one million points. Crucially, the method preserves exact GP uncertainty quantification without approximation of the posterior distribution. Empirical evaluation demonstrates that our approach achieves several-fold speedups over state-of-the-art sparse and variational GP methods while delivering superior predictive accuracy and calibrated uncertainty estimates.
Gaussian processes (GPs) face a dual challenge in large-scale model selection: prohibitive computational complexity—cubic time and quadratic memory in dataset size $N$—and the difficulty of preserving rigorous uncertainty quantification under approximation. This paper proposes a computation-aware GP model selection framework that explicitly models computational uncertainty and incorporates it into the hyperparameter optimization objective via a novel uncertainty-calibrated loss function. Integrating linear-time hyperparameter optimization with scalable approximate inference, our method achieves $O(N)$ time complexity while strictly maintaining calibrated predictive uncertainty. On a benchmark with 1.8 million data points, our approach completes end-to-end model selection within hours on a single GPU, outperforming state-of-the-art scalable GPs—including SGPR, CGGP, and SVGP—in both accuracy and uncertainty calibration. To the best of our knowledge, this is the first solution for large-scale GP deployment that simultaneously ensures theoretical rigor in uncertainty quantification and practical engineering feasibility.
Gaussian processes (GPs) face a computational bottleneck in large-scale regression due to $O(n^3)$ training complexity; while computation-aware GPs (CAGPs) reduce cost, their approximations often yield overly conservative uncertainty quantification. This paper introduces CAGP-GS—the first statistically calibrated CAGP framework. It establishes, for the first time, a theoretical equivalence between the calibration of probabilistic linear solvers and the calibration of CAGP posterior uncertainties. Leveraging a Gauss–Seidel–type probabilistic iterative solver, CAGP-GS jointly models computational uncertainty and explicitly calibrates the posterior distribution. The statistical calibration is rigorously validated on synthetic benchmarks. On a large-scale global temperature regression task, CAGP-GS achieves a superior trade-off between predictive reliability and computational efficiency—outperforming existing CAGPs in both uncertainty calibration and wall-clock time.
To address the computational bottleneck in Bayesian inference arising from repeated evaluations of expensive black-box models, this paper introduces “post-hoc Bayesian inference”: a new paradigm that constructs posterior approximations solely from existing model evaluations—such as those collected along a maximum a posteriori (MAP) optimization trajectory—without any additional model queries. The core method, variational sparse Bayesian quadrature (VSBQ), is the first to unify sparse Gaussian process modeling, Bayesian quadrature, and variational inference, enabling efficient posterior estimation under noisy or black-box likelihoods. Evaluated on synthetic benchmarks and a real-world computational neuroscience application, VSBQ achieves posterior accuracy comparable to standard MCMC or variational inference (VI) methods, while accelerating inference by one to two orders of magnitude—all using only pre-existing optimization trajectory data. This significantly reduces the computational cost of Bayesian inference without sacrificing fidelity.
Gaussian process regression (GPR) is often treated as a black-box surrogate model, limiting its pedagogical utility and interpretability in uncertainty quantification (UQ) for beginners. Core UQ tasks—including uncertainty propagation, risk estimation, Bayesian optimization, parameter inference, and sensitivity analysis—require deeper engagement with GPR’s inherent probabilistic structure. Method: This work develops a systematic, pedagogically grounded GPR-based UQ framework that integrates UQ-specific techniques: Bayesian quadrature, active learning, and surrogate-based sensitivity analysis. It emphasizes principled covariance kernel design, Bayesian hyperparameter estimation, and reproducible implementation. Contribution/Results: The framework lowers the barrier to applying GPR in complex UQ scenarios, enhances model transparency and decision reliability, and provides a theoretically rigorous yet practically actionable paradigm for UQ education and research across engineering and scientific disciplines.
This work addresses the challenge of applying Gaussian processes (GPs) to high-dimensional inputs, where they are prone to the curse of dimensionality, and existing two-stage dimensionality reduction approaches often compromise either predictive accuracy or reliable uncertainty quantification. To overcome this limitation, the authors propose an end-to-end Bayesian joint modeling framework that seamlessly integrates input dimensionality reduction within the GP formulation. By placing a prior on the Stiefel manifold to enforce orthogonality of the projection matrix and employing Riemannian Hamiltonian Monte Carlo for posterior inference along geodesics, the method achieves, for the first time, a unified Bayesian treatment of GP regression and dimensionality reduction. The framework is further extended to deep Gaussian processes to enhance representational capacity. Experimental results demonstrate superior performance over conventional two-stage methods in both prediction accuracy and uncertainty calibration, albeit at increased computational cost.
To address the computational inefficiency of solving large-scale linear systems in Gaussian process (GP) inference under incremental learning, this paper introduces warm-starting to iterative GP inference for the first time—leveraging historical small-scale posterior solutions as initial guesses for current large-scale iterative solvers. Exploiting the nested structure among sequential posteriors, the approach is compatible with diverse iterative linear solvers, including conjugate gradient, stochastic gradient descent, and alternating projections, thereby significantly accelerating convergence. Theoretical analysis demonstrates reduced iteration complexity, while empirical evaluation shows speedups of several-fold at equivalent accuracy. Moreover, in Bayesian optimization, the method achieves superior optimization performance under fixed computational budgets. This work establishes a novel, scalable paradigm for GP inference.
This work addresses the computational bottleneck of large-scale Gaussian processes, whose O(n³) complexity hinders efficient prediction. The authors propose a conditioning strategy based on carefully constructed data contrasts that exploits the low-rank structure of covariance matrices induced by smooth kernels. Within connected domains, the full conditional distribution can be accurately approximated using only a small number of linear combinations. This approach achieves O(T r²) offline precomputation and O(1) online prediction complexity at arbitrary locations, enabling real-time responses to unseen query points. Remarkably, the method attains near-linear overall computational cost while preserving machine-precision accuracy.
This study addresses the miscalibration of predictive uncertainty in Gaussian processes (GPs) arising from model misspecification by proposing the Prediction-Oriented Gaussian Process (PrO-GP). Moving beyond the limitations of conventional nonparametric Bayesian inference, this method explicitly treats the calibration of the predictive distribution as its core optimization objective. By leveraging a dimensionality reduction formulation coupled with a Markov chain Monte Carlo (MCMC) sampling algorithm, PrO-GP enables efficient, prediction-oriented Bayesian inference. Experiments conducted on both synthetic and real-world datasets demonstrate that PrO-GP substantially outperforms standard GPs, significantly improving the calibration accuracy of predictive uncertainty under model misspecification. These results establish PrO-GP as an effective framework for robust prediction in scenarios where standard GP assumptions are violated.
Existing Gaussian process (GP) methods struggle to simultaneously achieve computational efficiency, predictive accuracy, and model customizability on large-scale datasets—primarily due to reliance on approximations that compromise uncertainty quantification, restrict kernel and noise model design, and hinder effective modeling of expressive nonstationary patterns. This paper introduces gp2Scale: a framework that automatically constructs sparse covariance structures via flexible, compactly supported nonstationary kernels, enabling exact GP inference on datasets with millions of observations without any approximation. By integrating distributed computing with advanced sparse linear algebra techniques, gp2Scale efficiently solves linear systems and computes log-determinants. The method supports arbitrary kernel functions, noise models, and input space types—including irregular or high-dimensional domains. Empirical evaluation across multiple real-world benchmarks demonstrates that gp2Scale significantly outperforms or matches state-of-the-art approximate GP methods, thereby overcoming a fundamental bottleneck in scalable exact GP modeling.