surrogate modeling

Designs and builds surrogate (metamodel) systems that approximate expensive or high‑fidelity input→output mappings — including neural, kriging/Gaussian‑process, embedding‑based, functional, programmatic, geometric, local, and parameterized surrogate representations — and implements their construction and training, cold‑start and multi‑fidelity data fusion, and uncertainty quantification. Develops surrogate evaluation metrics and performance models and enforces constraint‑preserving or physics‑informed/regulatory behavior so the surrogate can accelerate design‑space evaluation, acquisition and optimization, trade‑off analysis, and closed‑loop workflows.

surrogatemodeling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.1
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$203K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of constructing efficient surrogate models for parametrized systems in multi-query scenarios—such as optimization, control, and uncertainty quantification—by proposing a unified scientific machine learning framework that systematically integrates physics-driven, data-driven, and hybrid modeling paradigms. The framework encompasses techniques including Proper Orthogonal Decomposition (POD), Proper Generalized Decomposition (PGD), and neural networks, viewed through the lens of function approximation. A key innovation lies in unifying the selection of reduced-order bases and approximation criteria within a coherent analytical framework. Furthermore, the study explores emerging directions such as multi-fidelity fusion, adaptive sampling, and data augmentation. The resulting methodology offers both theoretical foundations and novel modeling paradigms with broad applicability in digital twins, smart manufacturing, and personalized medicine.

data-drivenmulti-queryparametric systems

Bayesian Surrogate Training on Multiple Data Sources: A Hybrid Modeling Strategy

Dec 16, 2024
PR
Philipp Reiser
🏛️ University of Stuttgart | TU Dortmund University

To address the model mismatch arising from oversimplified simulation models and the underutilization of empirical measurements, this paper proposes a novel Bayesian surrogate modeling paradigm that jointly leverages simulation and real-world data. Our method introduces a dual-path, multi-source data fusion framework: (1) parallel posterior distribution ensembling and (2) end-to-end joint training, with the first explicit incorporation of empirical data into the Bayesian inference pipeline—endowing simulation models with diagnostic capability. The approach integrates Gaussian process regression, probabilistic distribution fusion, and uncertainty calibration. Evaluated on synthetic and real-world case studies, it achieves an average 23% reduction in RMSE, attains near-theoretical 95% credible interval coverage, and successfully detects structural deficiencies—including missing boundary conditions and omitted physical processes.

Address simulation model limitations using real-world measurement hintsDevelop probabilistic methods to integrate simulation and measurement dataImprove surrogate model accuracy by combining multiple data sources

This work addresses the challenge of constructing high-accuracy surrogate models in scenarios where high-fidelity data are scarce. We propose a novel multi-fidelity Gaussian process regression method that innovatively embeds low-fidelity data as augmented features into an expanded input space, thereby synergistically combining the strengths of co-kriging and autoregressive modeling. The approach achieves a balanced trade-off between modeling accuracy and computational efficiency within a unified framework, effectively leveraging heterogeneous multi-source data without requiring additional assumptions. Experimental results across multiple benchmark problems demonstrate that the proposed method significantly outperforms existing techniques, delivering higher predictive accuracy at lower computational cost.

Gaussian process regressionmachine learningmultifidelity

Two-stage Design for Failure Probability Estimation with Gaussian Process Surrogates

Oct 06, 2024
AS
Annie S. Booth
🏛️ Virginia Tech | Penn State

This work addresses the challenge of estimating small failure probabilities under stochastic inputs in computationally expensive deterministic simulations. We propose a two-stage adaptive budget allocation framework: in Stage I, a Gaussian process surrogate is sequentially trained using a contour-localization strategy; in Stage II, remaining simulation budget is greedily allocated to critical regions—guided by classification entropy—to perform high-fidelity evaluations. A hybrid Monte Carlo estimator is then constructed by integrating surrogate predictions with observed high-fidelity responses. Our method introduces the first “exploration–exploitation decoupled” budget allocation paradigm, overcoming reliability limitations inherent in pure surrogate-based Monte Carlo and importance sampling. Experiments across multiple benchmark functions and an airfoil flow simulation demonstrate that the approach achieves significantly improved accuracy and robustness using only several hundred high-fidelity evaluations.

Estimating failure probabilities with limited computational budgetImproving efficiency over existing sequential contour location methodsOptimizing surrogate model training for accurate classification

Surrogate-Based Optimization of System Architectures Subject to Hidden Constraints

Jul 27, 2024
JB
J. Bussemaker
🏛️ DLR | ONERA | Université de Toulouse

This work addresses the challenge of implicit constraints—manifested as evaluation failures—arising from unreliable physics-based simulations in system architecture optimization. To tackle this, we propose a surrogate modeling framework that integrates probabilistic feasibility prediction with Bayesian optimization. Methodologically, we introduce a novel hybrid discrete Gaussian process to model the Probability of Validity (PoV), coupled with an interior-point selection strategy based on a minimum PoV threshold; the framework natively supports hierarchical design variables and multi-objective optimization. Our approach achieves the first successful solution for a jet engine architecture optimization task with a 50% simulation failure rate. Across multiple synthetic benchmarks and real-world case studies, it significantly improves convergence robustness and optimization success rate. The implementation is publicly available as the SBArchOpt Python library.

Handling expensive and failed evaluations in optimizationOptimizing system architectures with hidden constraintsPredicting and managing failure regions in Bayesian Optimization

Latest Papers

What's happening recently
View more

This study addresses the high computational cost of high-fidelity modeling in composite materials, which arises from their multiscale nature, anisotropy, and coupling with manufacturing history, thereby hindering efficient design space exploration. To overcome this challenge, the work proposes a unified multifidelity surrogate modeling paradigm that systematically integrates co-kriging, autoregressive Gaussian processes, multifidelity deep Gaussian processes, and multifidelity neural networks. A general analytical framework is developed to model cross-fidelity correlations, characterize approximation errors, and quantify uncertainties. By effectively fusing abundant low-fidelity data with scarce high-fidelity observations, the proposed approach significantly enhances both predictive accuracy and computational efficiency. The methodology has been successfully demonstrated in forward design, inverse parameter identification, and heterogeneous data-driven engineering optimization workflows for composite materials.

composite materialsdesign space explorationhigh-fidelity simulation

This study addresses the challenges of high data requirements and ineffective fusion of multi-source, heterogeneous (multi-fidelity) data in agent-based modeling for manufacturing systems. The authors propose a hierarchical multi-task, multi-fidelity Gaussian process framework that decomposes each task’s response into a shared global trend and task-specific local residuals. By jointly modeling inter-task similarities and fidelity-level relationships, the approach enables efficient data fusion and rigorous uncertainty quantification. Notably, this work presents the first unified integration of multi-task learning and multi-fidelity modeling, accommodating an arbitrary number of tasks, design points, and fidelity levels. In both synthetic benchmarks and a real-world engine surface topography prediction case, the method achieves up to 19% and 23% higher prediction accuracy, respectively, compared to state-of-the-art multi-task models and independent stochastic kriging approaches.

data efficiencydata heterogeneitymulti-fidelity modeling

This study addresses the challenge of accurately capturing tail characteristics—corresponding to extreme samples—in solution field distributions when using neural network surrogates for uncertainty propagation. Taking the heat conduction equation as a benchmark, the work systematically evaluates the modeling capabilities of fully connected networks and DeepONets under both data-driven and physics-informed loss formulations, including weak-form residuals, with particular emphasis on tail prediction performance. The authors propose a method to identify extrapolative samples and find that fully connected networks trained with weak-form residual losses achieve superior accuracy under extreme inputs. Experimental results demonstrate that worst-case prediction errors for tail samples exceed those of the mean field by an order of magnitude, yet the proposed approach significantly enhances predictive accuracy for extreme scenarios on numerical datasets.

distribution tailsextreme samplesneural network surrogate models

This work addresses the challenge of high computational cost associated with high-fidelity simulations, which hinders extensive evaluation. To overcome this limitation, the authors propose a multi-fidelity surrogate modeling framework that integrates data of varying fidelity levels. The approach employs an ensemble of hierarchical Kriging models as base learners, whose predictions are combined via Bayesian model averaging. Crucially, the method introduces an innovative uncertainty quantification mechanism based on inter-model variance, which informs an adaptive sampling strategy to optimize the selection of training samples. Evaluated on multiple benchmark problems, the proposed framework consistently outperforms single-model approaches, achieving superior prediction accuracy, enhanced robustness, and improved data efficiency under constrained computational budgets.

adaptive samplingcomputational costheterogeneous information integration

Hot Scholars

PS

Paul Saves

IRIT, Université Toulouse Capitole
Machine LearningArtificial IntelligenceComputer ScienceApplied Mathematics
JM

Joseph Morlier

ISAE-SUPAERO and ICA-CNRS
multidisciplinary design optimizationtopology optimizationsurrogate modelingeco-informed material optimization
JB

Johannes Brandstetter

Johannes Kepler University (JKU) Linz
Deep LearningAI4ScienceAI4SimulationPhysics
ZM

Zeyuan Ma

South China University of Technology
Meta-Black-Box OptimizationReinforcement LearningLearning to Optimize
FA

Faez Ahmed

Associate Professor, MIT
Generative AIEngineering DesignMachine LearningEngineering Optimization