A Closer Look at Agentic BBO: Benchmarking LLM Agents for Black-Box Optimization

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the incomparability of evaluations and the difficulty of isolating design factors in existing research on black-box optimization (BBO) with LLM agents, which stem from variations in task domains and configurations. To this end, we propose the first cross-domain unified evaluation protocol and construct the AgenticBBO-Bench benchmark. Under a unified budget, we systematically evaluate agent performance across synthetic functions and real-world scenarios, providing an in-depth analysis of how tools, prior knowledge, and LLM roles influence outcomes. Experimental results demonstrate that agent-based methods outperform direct LLM invocation in five domains and surpass the best numerical optimizers in four. Furthermore, models such as GPT-6 Astra are shown to lie on the performance-cost Pareto frontier.
📝 Abstract
Black-box optimization (BBO) arises in many scientific and engineering problems where objective evaluations are expensive and limited. Recent large language model (LLM) agents offer a new way to approach BBO by combining task semantics, computation, optimization tools, and feedback-driven decision making, showing great potential due to the integration with mathematically rigorous tools. However, existing agentic BBO studies use different task domains and system configurations, making their results difficult to compare and the effects of individual design choices hard to isolate. We therefore introduce AgenticBBO-Bench, a cross-domain benchmark for agentic BBO spanning synthetic functions, hyperparameter optimization, database tuning, chip design, and molecular design under a unified finite-budget evaluation protocol. In our experiments, agentic BBO achieves higher family-averaged scores than direct LLM-based methods in all five domains and outperforms the best numerical optimizers in four. We further study three factors shaping agent performance: optimization tools, task information and prior knowledge, and the role of the LLM during search. Our results show that additional numerical tools do not consistently improve performance, task semantics are broadly useful while more specific priors are less reliable, and numerical optimizers can effectively absorb gains from search trajectories established by the agent. Finally, we introduce a five-task frontier challenge within AgenticBBO-Bench and evaluate seven LLMs under the Codex agent harness, where GPT-6 Astra and DeepSeek-V4.1-Flash lie on the Pareto frontier of performance and cost among the evaluated models. Our code is available at https://github.com/lamda-bbo/agentic-bbo.
Problem

Research questions and friction points this paper is trying to address.

Black-Box Optimization
LLM Agents
Benchmarking
Agentic BBO
Innovation

Methods, ideas, or system contributions that make the work stand out.

Black-Box Optimization
LLM Agents
Benchmark
Cross-Domain Evaluation
Pareto Frontier
🔎 Similar Papers
Ming Chen
Ming Chen
State Key Laboratory of Novel Software Technology, Nanjing University; School of Artificial Intelligence, Nanjing University
Rong-Xi Tan
Rong-Xi Tan
Nanjing University
Black-box optimizationLearning to Optimize
Ke Xue
Ke Xue
Nanjing University
Black-Box OptimizationMachine Learning
Y
Yu-Jie Zhou
State Key Laboratory of Novel Software Technology, Nanjing University; School of Artificial Intelligence, Nanjing University
T
Taiye Lu
State Key Laboratory of Novel Software Technology, Nanjing University; School of Artificial Intelligence, Nanjing University
Z
Zhi-Xuan Gao
State Key Laboratory of Novel Software Technology, Nanjing University; School of Artificial Intelligence, Nanjing University
P
Peng Xie
State Key Laboratory of Novel Software Technology, Nanjing University; School of Artificial Intelligence, Nanjing University
Z
Zijun Shen
State Key Laboratory of Novel Software Technology, Nanjing University; School of Artificial Intelligence, Nanjing University
C
Chen Lu
State Key Laboratory of Novel Software Technology, Nanjing University; School of Artificial Intelligence, Nanjing University
H
Haopu Shang
State Key Laboratory of Novel Software Technology, Nanjing University; School of Artificial Intelligence, Nanjing University
Chao Qian
Chao Qian
Nanjing University
Artificial intelligenceevolutionary algorithmsmachine learning