gaussian process surrogate

Designs and validates Gaussian process–based surrogate models (data-driven emulators) that approximate the outputs of expensive simulations or experiments from input variables while providing calibrated probabilistic uncertainty estimates. These data-driven emulators — including application-specific surrogate engine models — are built to generate realistic synthetic trajectories, enable fast, repeatable evaluations for controllers and testing, and support controlled analysis and scenario generation.

gaussianprocesssurrogate

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.18
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Fast Emulation and Modular Calibration for Simulators with Functional Response

May 25, 2024
GH
Grant Hutchings
🏛️ Los Alamos National Laboratory | Simon Fraser University

To address the scalability challenge in constructing surrogate models for functional-response computer models—such as those producing spatiotemporal or time-series outputs—this paper proposes a highly scalable hybrid framework. The method integrates global input-space adaptive length-scale scaling with local approximate Gaussian processes (Gramacy–Apley type), preserving the modeling fidelity of established functional-response approaches (e.g., Higdon and Kennedy–O’Hagan) while overcoming the cubic computational bottleneck of standard Gaussian processes. Designed for modular calibration, it is particularly suited to multi-physics fluid dynamics simulators like FLAG. Evaluated on a dataset of 20,000 FLAG simulations, the framework achieves order-of-magnitude speedup in prediction time without sacrificing accuracy. The implementation is publicly available as the R package FlaGP.

Fast calibration of multiphysics simulators with large ensemblesHandling dense functional output like spatial or time-series dataScalable surrogate models for efficient computer simulator emulation

Bayesian Surrogate Training on Multiple Data Sources: A Hybrid Modeling Strategy

Dec 16, 2024
PR
Philipp Reiser
🏛️ University of Stuttgart | TU Dortmund University

To address the model mismatch arising from oversimplified simulation models and the underutilization of empirical measurements, this paper proposes a novel Bayesian surrogate modeling paradigm that jointly leverages simulation and real-world data. Our method introduces a dual-path, multi-source data fusion framework: (1) parallel posterior distribution ensembling and (2) end-to-end joint training, with the first explicit incorporation of empirical data into the Bayesian inference pipeline—endowing simulation models with diagnostic capability. The approach integrates Gaussian process regression, probabilistic distribution fusion, and uncertainty calibration. Evaluated on synthetic and real-world case studies, it achieves an average 23% reduction in RMSE, attains near-theoretical 95% credible interval coverage, and successfully detects structural deficiencies—including missing boundary conditions and omitted physical processes.

Address simulation model limitations using real-world measurement hintsDevelop probabilistic methods to integrate simulation and measurement dataImprove surrogate model accuracy by combining multiple data sources

Robust designs for Gaussian process emulation of computer experiments

Jul 12, 2025
SM
Simon Mak
🏛️ Duke University | Georgia Institute of Technology

To address the sensitivity of experimental designs to response surface smoothness and their computational inefficiency in high-dimensional Gaussian process surrogate modeling, this paper proposes two robust design classes: support points and projected support points. These designs are efficiently constructed in high dimensions and for large sample sizes via difference-of-convex programming (DCP), ensuring both theoretical interpretability and computational scalability. Theoretical analysis establishes their superior properties in terms of space-fillingness, uniformity, and stability. Numerical experiments demonstrate that the proposed designs consistently outperform classical alternatives—including Latin hypercube sampling (LHS) and maximin designs—across smooth, rough, and highly oscillatory response surfaces. Specifically, they yield 15–30% improvements in model prediction accuracy and reduce generalization error variance by approximately 40%. This work provides a unified design framework and practical algorithms for robust, adaptive computer experiment modeling.

Efficient generation of large high-dimensional designsGood performance across smooth and rugged surfacesRobust experimental designs for Gaussian process emulation

This work addresses the challenge of performing Bayesian inference in high-computational-cost settings, where standard approaches are often infeasible and existing surrogate modeling methods frequently neglect uncertainty propagation and workflow integration. The paper introduces the first unified framework that systematically integrates surrogate modeling, Bayesian inference, uncertainty quantification, and active learning. By co-designing surrogate uncertainty modeling with sequential optimization mechanisms, the framework enables an efficient and robust inference pipeline. It synthesizes previously fragmented research across multiple domains into a coherent methodology, offering practitioners and method developers a clear, actionable surrogate-based Bayesian workflow that substantially enhances both the reliability and efficiency of inference in expensive simulation scenarios.

Active learningBayesian inferenceEmulators

Two-stage Design for Failure Probability Estimation with Gaussian Process Surrogates

Oct 06, 2024
AS
Annie S. Booth
🏛️ Virginia Tech | Penn State

This work addresses the challenge of estimating small failure probabilities under stochastic inputs in computationally expensive deterministic simulations. We propose a two-stage adaptive budget allocation framework: in Stage I, a Gaussian process surrogate is sequentially trained using a contour-localization strategy; in Stage II, remaining simulation budget is greedily allocated to critical regions—guided by classification entropy—to perform high-fidelity evaluations. A hybrid Monte Carlo estimator is then constructed by integrating surrogate predictions with observed high-fidelity responses. Our method introduces the first “exploration–exploitation decoupled” budget allocation paradigm, overcoming reliability limitations inherent in pure surrogate-based Monte Carlo and importance sampling. Experiments across multiple benchmark functions and an airfoil flow simulation demonstrate that the approach achieves significantly improved accuracy and robustness using only several hundred high-fidelity evaluations.

Estimating failure probabilities with limited computational budgetImproving efficiency over existing sequential contour location methodsOptimizing surrogate model training for accurate classification

Latest Papers

What's happening recently
View more

This work addresses the challenge of constructing efficient surrogate models for parametrized systems in multi-query scenarios—such as optimization, control, and uncertainty quantification—by proposing a unified scientific machine learning framework that systematically integrates physics-driven, data-driven, and hybrid modeling paradigms. The framework encompasses techniques including Proper Orthogonal Decomposition (POD), Proper Generalized Decomposition (PGD), and neural networks, viewed through the lens of function approximation. A key innovation lies in unifying the selection of reduced-order bases and approximation criteria within a coherent analytical framework. Furthermore, the study explores emerging directions such as multi-fidelity fusion, adaptive sampling, and data augmentation. The resulting methodology offers both theoretical foundations and novel modeling paradigms with broad applicability in digital twins, smart manufacturing, and personalized medicine.

data-drivenmulti-queryparametric systems

Traditional Gaussian processes struggle to model stochastic simulators with discrete, heteroscedastic, or non-Gaussian outputs. This work proposes a scalable Generalized Deep Gaussian Process (GDGP) framework that unifies the treatment of diverse non-Gaussian responses—such as Poisson, negative binomial, and categorical—through a latent Gaussian process structure. By integrating the Vecchia approximation, the approach enables efficient Bayesian inference for large-scale inputs and replicated simulations. GDGP represents the first extension of deep Gaussian processes to general non-Gaussian output settings, substantially enhancing the ability to model complex, non-stationary simulation systems. Empirical evaluations on both synthetic and real-world case studies demonstrate the superior performance of GDGP, and the accompanying R package dgpsi is publicly released to facilitate broader application.

discrete outputsGaussian process emulationheteroskedasticity

Efficient multi-fidelity Gaussian process regression for noisy outputs and non-nested experimental designs

Nov 25, 2025
NB
Nils Baillie
🏛️ Université Paris-Saclay | CEA | École polytechnique | Institut Polytechnique de Paris | CMAP | CNRS

This paper addresses Gaussian process regression modeling with non-nested, noisy multi-fidelity data. We propose an efficient, scalable multi-fidelity surrogate model that abandons the conventional recursive autoregressive assumption and instead introduces parameterized linear predictors, for which we derive closed-form update rules and integrate an expectation-maximization (EM) algorithm to estimate high-fidelity hyperparameters. A decoupled optimization strategy is further designed to substantially reduce computational complexity. Our key contributions are threefold: (i) the first method supporting simultaneous non-nested data structures and observation noise in multi-source fusion; (ii) a theoretically grounded analytical framework for learning closed-form solutions; and (iii) consistent superiority over state-of-the-art approaches in both prediction accuracy and training efficiency across multiple benchmarks and real-world tasks, with systematic experiments confirming strong generalizability and scalability.

Enables efficient parameter estimation via EM algorithm optimizationGeneralizes multi-fidelity Gaussian process for noisy non-nested dataProvides scalable surrogate modeling for complex multi-fidelity applications

This work addresses the limitations of conventional Gaussian processes in effectively modeling multi-output stochastic simulators and explicitly disentangling regression trends from residual variations. To overcome these challenges, the authors propose the Multi-Output Orthogonal Gaussian Process (MOOGP), which introduces orthogonality constraints into the covariance function to explicitly decouple trend and residual components, thereby enforcing an orthogonal covariance structure. A structured likelihood formulation is further developed to enhance computational efficiency. Experimental results demonstrate that MOOGP accurately recovers true underlying trends—avoiding sign reversals—and achieves superior predictive performance and interpretability in complex scenarios such as heavy-ion collision simulations.

Gaussian processesmulti-outputnoisy simulators

This study addresses the challenge of predicting bending deformation in pressurized water reactor fuel assemblies, a problem governed by strongly coupled fluid–structure interactions that render full-order multiphysics simulations computationally prohibitive for large-scale uncertainty quantification. To overcome this, the authors propose a Gaussian process surrogate modeling framework that replaces expensive high-fidelity simulations, enabling efficient and rigorous uncertainty propagation and global sensitivity analysis. The key innovation lies in the development of the first theoretical framework guaranteeing bounded predictive variance for coupled Gaussian processes, thereby preventing uncertainty divergence during iterative refinement and establishing a rigorous mathematical foundation for reliable surrogate use in multiphysics systems. Numerical experiments demonstrate that the method achieves high prediction accuracy and stability while substantially reducing computational cost, confirming its effectiveness and scalability.

computational costfluid-structure interactionfuel assembly bow

Hot Scholars

KL

Keqiang Li

Department of Automotive Engineering, Tsinghua University
Intelligent VehiclesAdvanced Driver Assistant Systems
AV

Aki Vehtari

Professor, Aalto University
Bayesian analysisBayesian statisticsGaussian processesProbabilistic programming
KK

Kyurae Kim

PhD Student, University of Pennsylvania
Bayesian inferencestochastic optimizationmachine learningsignal processing