Beyond Scripted Search: Sample-Efficient Reward Discovery via Agentic Black-box Optimization

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of designing dense rewards in reinforcement learning and the low sample efficiency of scripted search methods by proposing the ARBO framework. Departing from fixed prompt-generation paradigms, this approach leverages large language model agents to autonomously maintain belief states and invoke tools at runtime, dynamically constructing candidate reward functions based on historical evaluations to achieve efficient black-box optimization. Furthermore, a persistent workspace mechanism is introduced to support continuous policy evolution. Experimental results demonstrate that, under equivalent computational budgets, ARBO improves manipulation success rates by 29.9% and power grid task scores by 192.8%, significantly enhancing both sample efficiency and generalization performance.
📝 Abstract
Designing dense reward functions for low-level reinforcement learning (RL) control remains difficult. Recent work uses large language models (LLMs) to iteratively generate and refine reward functions using policy-training feedback within scripted search algorithms. However, evaluating each candidate requires a full RL training run, making sample efficiency a central challenge for reward search on complex control tasks. To address this limitation, we propose an Agentic Reward Black-box Optimization (ARBO) framework, in which an LLM agent builds the search strategy at run time from an evaluation history maintained as its persistent workspace. The evaluation history comprises two components: observations maintained by the evaluation oracle, including candidate scores, per-term training curves, and error tracebacks; and an agent-maintained belief that records diagnoses and intended next steps. The agent queries both with tools and generates the next batch of reward candidates, rather than generating them in a single pass from a fixed prompt. Across four control domains, ARBO achieves gains of 29.9% in manipulation success rate and 192.8% in power-grid score over baseline means under a shared evaluation budget. Ablations examine each component's contribution and sensitivity to backbone choice.
Problem

Research questions and friction points this paper is trying to address.

reward function design
reinforcement learning
sample efficiency
black-box optimization
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic Black-box Optimization
Reward Discovery
Large Language Models
Sample Efficiency
Reinforcement Learning