How Should a Prompt Optimizer Spend a Tight Budget? BudgetAPO with Noise-Adaptive Evaluation

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of existing Automatic Prompt Optimization (APO) methods to resource exhaustion and noise interference under strict invocation budgets. We propose BudgetAPO, a single-stage optimization framework that introduces a short-probe-based noise-adaptive slicing mechanism, integrated with paired comparative statistical testing and a reflective joint rewriting strategy, to achieve efficient prompt optimization within limited budgets. Experiments demonstrate that BudgetAPO attains state-of-the-art performance across seven benchmarks and five models. Notably, under a stringent constraint of only 250 invocations, it achieves a failure rate as low as 13%, significantly outperforming baselines such as GEPA. This work provides an efficient and reliable solution for prompt optimization in budget-constrained scenarios.
📝 Abstract
Automatic prompt optimization (APO) has been widely employed to adapt large language models without updating their weights, yielding promising results. However, existing methods such as GEPA and OPRO assume hundreds to thousands of subject-model calls, far more than is practical behind paid, rate-limited APIs. Under tight budgets they fail in two ways: multi-stage pipelines can exhaust the budget and return the seed prompt unchanged, while single-stage methods compare candidates on fixed-size minibatches, regardless of each task's noise. As a remedy, we introduce BudgetAPO, a single-stage optimizer for the tight-budget regime. BudgetAPO incorporates (1) a noise-adaptive rule that sizes the evaluation slice to each task's noise, measured by a short probe; (2) a fixed slice that turns every accept/reject decision into a paired comparison; and (3) a reflective operator that rewrites reasoning strategy and output format jointly. Extensive results across seven benchmarks and five subject models demonstrate that BudgetAPO ranks first on every subject and beats every baseline under Holm-corrected paired tests, while returning the seed in 13% of runs at 250 calls against 86% for GEPA. On GPT-OSS-20B, GEPA needs 4.5 times as many calls to match BudgetAPO's 100-call score.
Problem

Research questions and friction points this paper is trying to address.

Automatic Prompt Optimization
Tight Budget
Noise-Adaptive Evaluation
Large Language Models
API Rate Limit
Innovation

Methods, ideas, or system contributions that make the work stand out.

Automatic Prompt Optimization
Noise-Adaptive Evaluation
Budget-Constrained Optimization
Paired Comparison
Reflective Operator
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Haoyue Liu
Haoyue Liu
School of Artificial Intelligence and Automation, Huazhong University of Science and Technology
Computer VisionEvent Camera
Zhichao Wang
Zhichao Wang
School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen
H
Huanyu Yan
School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen
X
Xiaoying Tang
School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen; Shenzhen Future Network of Intelligence Institute (FNii-Shenzhen); Guangdong Provincial Key Laboratory of Future Networks of Intelligence, CUHK(SZ)