Optimizing the Unknown: Black Box Bayesian Optimization with Energy-Based Model and Reinforcement Learning

📅 2025-10-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Standard Bayesian optimization (BO) suffers from a “one-step bias” due to its greedy, single-step acquisition strategy, leading to premature convergence to local optima—particularly detrimental in high-dimensional, complex black-box optimization. To address this, we propose REBMBO, a reinforcement learning–enhanced BO framework that formulates BO as a Markov decision process. REBMBO jointly models local fidelity and global structure by integrating Gaussian processes (GPs) with energy-based models, and employs proximal policy optimization (PPO) to enable adaptive, multi-step lookahead balancing exploration and exploitation. Its key contribution is the first application of reinforcement learning to learn multi-step BO policies, dynamically adjusting both exploration depth and direction. Experiments on synthetic benchmarks and real-world tasks demonstrate that REBMBO consistently outperforms state-of-the-art BO methods, exhibiting strong robustness and generalization across diverse GP kernels.

Technology Category

Search and Optimization: Sampling/Simulation-based SearchIntelligent Robots: Learning & Optimization for ROBMachine Learning: Reinforcement Learning

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingEconomics, Online Markets and Human Computation: Uses of LLMs and GenAI for marketplace design, bidding, and strategic interactions
📝 Abstract
Existing Bayesian Optimization (BO) methods typically balance exploration and exploitation to optimize costly objective functions. However, these methods often suffer from a significant one-step bias, which may lead to convergence towards local optima and poor performance in complex or high-dimensional tasks. Recently, Black-Box Optimization (BBO) has achieved success across various scientific and engineering domains, particularly when function evaluations are costly and gradients are unavailable. Motivated by this, we propose the Reinforced Energy-Based Model for Bayesian Optimization (REBMBO), which integrates Gaussian Processes (GP) for local guidance with an Energy-Based Model (EBM) to capture global structural information. Notably, we define each Bayesian Optimization iteration as a Markov Decision Process (MDP) and use Proximal Policy Optimization (PPO) for adaptive multi-step lookahead, dynamically adjusting the depth and direction of exploration to effectively overcome the limitations of traditional BO methods. We conduct extensive experiments on synthetic and real-world benchmarks, confirming the superior performance of REBMBO. Additional analyses across various GP configurations further highlight its adaptability and robustness.
Problem

Research questions and friction points this paper is trying to address.

Overcoming one-step bias in Bayesian Optimization for complex tasks
Optimizing costly black-box functions without gradient information
Addressing local optima convergence in high-dimensional optimization problems
Innovation

Methods, ideas, or system contributions that make the work stand out.

Integrates Gaussian Processes with Energy-Based Model
Defines Bayesian Optimization as Markov Decision Process
Uses Proximal Policy Optimization for adaptive exploration
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Ruiyao Miao
University of California, Los Angeles
J
Junren Xiao
The Hong Kong University of Science and Technology (Guangzhou)
S
Shiya Tsang
The Hong Kong University of Science and Technology (Guangzhou)
Hui Xiong
Hui Xiong
Senior Scientist, Candela Corporation
Ultrafast dynamicsatomic molecular physicsfree electron laser
Y
Yingnian Wu
University of California, Los Angeles