🤖 AI Summary
This study addresses the vulnerability of existing Automatic Prompt Optimization (APO) methods to resource exhaustion and noise interference under strict invocation budgets. We propose BudgetAPO, a single-stage optimization framework that introduces a short-probe-based noise-adaptive slicing mechanism, integrated with paired comparative statistical testing and a reflective joint rewriting strategy, to achieve efficient prompt optimization within limited budgets. Experiments demonstrate that BudgetAPO attains state-of-the-art performance across seven benchmarks and five models. Notably, under a stringent constraint of only 250 invocations, it achieves a failure rate as low as 13%, significantly outperforming baselines such as GEPA. This work provides an efficient and reliable solution for prompt optimization in budget-constrained scenarios.
📝 Abstract
Automatic prompt optimization (APO) has been widely employed to adapt large language models without updating their weights, yielding promising results. However, existing methods such as GEPA and OPRO assume hundreds to thousands of subject-model calls, far more than is practical behind paid, rate-limited APIs. Under tight budgets they fail in two ways: multi-stage pipelines can exhaust the budget and return the seed prompt unchanged, while single-stage methods compare candidates on fixed-size minibatches, regardless of each task's noise. As a remedy, we introduce BudgetAPO, a single-stage optimizer for the tight-budget regime. BudgetAPO incorporates (1) a noise-adaptive rule that sizes the evaluation slice to each task's noise, measured by a short probe; (2) a fixed slice that turns every accept/reject decision into a paired comparison; and (3) a reflective operator that rewrites reasoning strategy and output format jointly. Extensive results across seven benchmarks and five subject models demonstrate that BudgetAPO ranks first on every subject and beats every baseline under Holm-corrected paired tests, while returning the seed in 13% of runs at 250 calls against 86% for GEPA. On GPT-OSS-20B, GEPA needs 4.5 times as many calls to match BudgetAPO's 100-call score.