🤖 AI Summary
To address Bayesian calibration of complex systems under computationally expensive, data-sparse, or noisy conditions, this paper proposes a sequential surrogate modeling framework. It employs Gaussian process regression to progressively approximate the log-likelihood and its gradients, integrated with gradient-driven Metropolis-adjusted Langevin algorithm (MALA) sampling. We introduce an uncertainty-aware adaptive likelihood evaluation strategy that invokes the expensive true likelihood only in regions where the surrogate exhibits high predictive uncertainty and posterior sensitivity. This enables tight coupling between surrogate modeling and gradient-informed MCMC. Evaluated on synthetic benchmarks and a high-speed train parameter calibration case, the method scales to >20-dimensional parameter spaces, achieves significantly accelerated convergence, reduces computational cost by over 60%, and accurately recovers physically meaningful posterior distributions—even under missing-sensor conditions.
📝 Abstract
Numerical simulations are crucial for modeling complex systems, but calibrating them becomes challenging when data are noisy or incomplete and likelihood evaluations are computationally expensive. Bayesian calibration offers an interesting way to handle uncertainty, yet computing the posterior distribution remains a major challenge under such conditions. To address this, we propose a sequential surrogate-based approach that incrementally improves the approximation of the log-likelihood using Gaussian Process Regression. Starting from limited evaluations, the surrogate and its gradient are refined step by step. At each iteration, new evaluations of the expensive likelihood are added only at informative locations, that is to say where the surrogate is most uncertain and where the potential impact on the posterior is greatest. The surrogate is then coupled with the Metropolis-Adjusted Langevin Algorithm, which uses gradient information to efficiently explore the posterior. This approach accelerates convergence, handles relatively high-dimensional settings, and keeps computational costs low. We demonstrate its effectiveness on both a synthetic benchmark and an industrial application involving the calibration of high-speed train parameters from incomplete sensor data.