🤖 AI Summary
Standard Bayesian optimization (BO) suffers from a “one-step bias” due to its greedy, single-step acquisition strategy, leading to premature convergence to local optima—particularly detrimental in high-dimensional, complex black-box optimization. To address this, we propose REBMBO, a reinforcement learning–enhanced BO framework that formulates BO as a Markov decision process. REBMBO jointly models local fidelity and global structure by integrating Gaussian processes (GPs) with energy-based models, and employs proximal policy optimization (PPO) to enable adaptive, multi-step lookahead balancing exploration and exploitation. Its key contribution is the first application of reinforcement learning to learn multi-step BO policies, dynamically adjusting both exploration depth and direction. Experiments on synthetic benchmarks and real-world tasks demonstrate that REBMBO consistently outperforms state-of-the-art BO methods, exhibiting strong robustness and generalization across diverse GP kernels.
📝 Abstract
Existing Bayesian Optimization (BO) methods typically balance exploration and exploitation to optimize costly objective functions. However, these methods often suffer from a significant one-step bias, which may lead to convergence towards local optima and poor performance in complex or high-dimensional tasks. Recently, Black-Box Optimization (BBO) has achieved success across various scientific and engineering domains, particularly when function evaluations are costly and gradients are unavailable. Motivated by this, we propose the Reinforced Energy-Based Model for Bayesian Optimization (REBMBO), which integrates Gaussian Processes (GP) for local guidance with an Energy-Based Model (EBM) to capture global structural information. Notably, we define each Bayesian Optimization iteration as a Markov Decision Process (MDP) and use Proximal Policy Optimization (PPO) for adaptive multi-step lookahead, dynamically adjusting the depth and direction of exploration to effectively overcome the limitations of traditional BO methods. We conduct extensive experiments on synthetic and real-world benchmarks, confirming the superior performance of REBMBO. Additional analyses across various GP configurations further highlight its adaptability and robustness.