🤖 AI Summary
This study addresses the cost-inefficiency of LLM-guided program evolution by proposing a cost-aware evolutionary framework designed to maximize performance gains per unit cost. Methodologically, it introduces a novel strong-weak model division-of-labor mechanism that synergizes exploration and code implementation, alongside a prefix-sharing caching strategy to optimize resource utilization. Furthermore, a budget-aware area under the curve (BA-AUC) metric is formulated for evaluation. Experimental results demonstrate that this framework effectively integrates evolutionary algorithms with prompt engineering, surpassing state-of-the-art baselines across multiple benchmarks at minimal expense. Notably, it establishes a new record on the circle packing task for merely $0.55, underscoring its exceptional cost-effectiveness.
📝 Abstract
LLM-guided evolutionary methods, such as AlphaEvolve, have emerged as powerful approaches for challenging computational optimization problems, such as circle packing. However, prior work typically optimizes performance gain over a fixed number of iterations. We argue that practical optimization should maximize gain per unit cost. To this end, we propose FrugalEvo, a cost-aware evolutionary framework where a stronger, higher-cost LLM explores solution strategies, and a cheaper LLM implements them and iteratively refines the resulting code. We also design a cache-efficient evolution process, where our harness and prompts maximize the sharing of prefixes across different evolution steps, to improve cache reuse. To measure solution quality throughout a fixed cost budget, we introduce Budget-Aware Area Under the Curve (BA-AUC), defined as the area under the best-so-far evaluation score curve over cumulative LLM cost, up to the budget. Across 10 mathematical and systems optimization tasks, FrugalEvo matches or surpasses state-of-the-art baselines, including OpenEvolve, ShinkaEvolve, AdaEvolve, and EvoX, in final solution quality and achieves higher BA-AUC on 9 tasks. It also achieves higher average performance than these baselines on 10 algorithmic optimization tasks from ALE-Bench-Lite. Notably, on circle packing, FrugalEvo achieves new state-of-the-art performance with GPT-5.6 Terra and Luna for only 1.68 USD and with GLM-5.3 and its Flash variant for only 0.55 USD, matching or surpassing all baselines, including multi-agent methods such as CORAL and SwarmResearch, which cost approximately 50 USD on average.