🤖 AI Summary
This study addresses the challenge that limited data in high-dimensional Bayesian optimization often causes automatic relevance determination (ARD) models to suffer from direction invisibility or insufficiently constrained priors. To overcome this, we propose the Iso-BO framework, which elucidates the dependence mechanism of the marginal log-likelihood on ARD length scales. Specifically, this work replaces the ARD kernel with an isotropic Gaussian process employing a shared length scale, thereby eliminating small-sample modeling deficiencies by removing coordinate reweighting while preserving the standard optimization pipeline. Extensive evaluations on both synthetic and real-world benchmarks demonstrate that the proposed method significantly outperforms mainstream Bayesian optimization algorithms. Furthermore, Iso-BO exhibits remarkable robustness even under strongly anisotropic scenarios, offering a principled and effective solution for sample-efficient high-dimensional optimization.
📝 Abstract
High-dimensional Bayesian optimization (BO) often fits Gaussian process (GP) surrogates from far fewer observations than input dimensions. Modern Vanilla BO can perform well in this regime with dimension-aware priors, initialization, and acquisition optimization, but it typically retains automatic relevance determination (ARD), fitting one lengthscale per input coordinate. We study this modeling choice and propose Iso-BO, a controlled modification that replaces the ARD GP with an isotropic GP using one shared lengthscale while keeping the surrounding BO pipeline matched. For radial kernels, we show that the marginal log likelihood (MLL) depends on the inverse-squared ARD lengthscales only through weighted pairwise distances among the observed inputs. The current design can therefore leave some ARD directions exactly invisible or only weakly constrained by the MLL. Iso-BO removes coordinatewise reweighting and fits a single shared scale instead. Lengthscale-fitting and predictive-density diagnostics show that this finite-data effect appears in practice, including when the data-generating process is anisotropic. Across GP-prior, synthetic, and real-world benchmarks, Iso-BO often improves over matched modern Vanilla BO and remains competitive with the included high-dimensional BO baselines under the tested budgets. Stress tests also show the expected boundary wherein sufficiently strong, learnable anisotropy can favor the more flexible ARD model.