🤖 AI Summary
This study addresses the low sample efficiency in expensive simulator inference and the underutilization of gradient information. We propose an active learning framework based on Bayesian optimization that integrates gradients obtained via automatic differentiation into Gaussian process surrogate models. Furthermore, this work presents the first systematic evaluation of the differential augmentation benefits between forward-mode and reverse-mode gradients under finite computational budgets. Experimental results demonstrate that reverse-mode gradients significantly accelerate convergence, with optimization gains sufficient to offset the additional computational overhead, whereas forward-mode gradients yield limited improvements. Overall, this research provides critical empirical evidence supporting the application of gradient-enhanced surrogate models for efficient simulation-based inference.
📝 Abstract
Simulators based on differential equations are ubiquitous in science and engineering. They are often used in simulation-based inference to evaluate the posterior distribution of the input parameters based on real-world observations of the simulator outputs. However, inference becomes challenging when individual simulator evaluations are computationally expensive. In such cases, a Bayesian optimization-based active learning approach with Gaussian process surrogate models has been used to maximize the information obtained from a limited simulation budget. Recently, gradients of simulator outputs with respect to input parameters have become increasingly available, yet they are rarely exploited for inference. Even though we only need to learn the simulator input-output relationship, gradient information can provide an additional valuable signal to guide the active learning procedure. This is of particular interest in the case of expensive simulators, when sample efficiency is crucial.
In this paper, we demonstrate how incorporating gradient information into the Gaussian process surrogate accelerates Bayesian optimization-based inference under a limited simulation budget. Our results show significant improvement in convergence speed from using gradient information. For reverse-mode differentiation, the inference efficiency gains are maintained when accounting for the additional computational cost. In contrast, for forward-mode differentiation, the inference speed-up does not outweigh the computational costs. These results indicate that gradient-enhanced surrogates are beneficial primarily in problems where the number of parameters exceeds the output dimensionality, where reverse-mode differentiation is efficient.