🤖 AI Summary
This study addresses the prohibitive computational cost of gradient observations in Gaussian process (GP) surrogate models by proposing a bidirectional budgeted derivative GP method. Through an innovative bidirectional gradient aggregation mechanism, this approach compresses md gradient coordinates per objective into at most 2m directional derivatives, separately handling direct predictive contributions and indirect correlations. By integrating Vecchia sparse approximations with directional derivative projections, it achieves O(m³) dense decomposition complexity and provides a theoretical analysis of posterior error. Experimental results demonstrate that the proposed method matches the accuracy of state-of-the-art exact approaches while significantly reducing resource consumption. It supports ultra-large conditioning sets and attains superior predictive accuracy at a lower computational cost than pure function-evaluation GP baselines.
📝 Abstract
Gradient observations promise more accurate Gaussian process (GP) surrogates, but the cost of incorporating them has long stood in the way of realizing that promise. We propose a derivative GP with a budget of just two directions per observed gradient. One direction focuses on each gradient's direct contribution to target prediction, while the other aggregates its indirect contributions through correlations with the conditioning function values. Within a Vecchia approximation, where each prediction conditions on $m$ nearby inputs in $d$ dimensions, this construction represents their $md$ gradient coordinates using at most $2m$ directional derivatives, giving $\mathcal{O}(m^3)$ dense factorization cost per prediction target. For general conditioning sets, we bound the posterior approximation error relative to using full gradients and characterize when the error is small or the approximation is exact. In simulations, our method matches the accuracy of a leading exact gradient-reduction method at equal conditioning set size. Because its cost grows much more slowly with that size, it can use conditioning sets well beyond the memory limit of the exact method, reaching lower prediction error with a small fraction of the time and memory. Notably, our method can exploit gradient observations while requiring less computation time or memory than function-only GP baselines.