π€ AI Summary
This study addresses the computational waste and inefficient screening of high-value candidates caused by fixed resource allocation in LLM program evolution. To overcome these limitations, this work proposes a self-evolving resource allocation agent that integrates large language models, program evolution, and reinforcement learning-based experience consolidation. The method adaptively revises allocation strategies by dynamically accumulating search experiences and incorporates counterfactual reasoning to optimize exploration efficiency. Experimental results on coding benchmarks demonstrate that the proposed framework reduces evaluation counts by 59β82% and token consumption by 61β89%, while achieving performance improvements of 8.7β12.0% under equivalent budgets. These findings confirm that the approach enables dynamic adaptation of resource allocation and delivers substantial efficiency gains in evolutionary program synthesis.
π Abstract
LLM-based program evolution relies on evaluation feedback to guide the iterative search for high-performing programs. However, evaluation is often computationally expensive, making it essential to allocate limited resources to candidates that can most effectively advance the search. Existing LLM-based methods typically rely on fixed allocation strategies throughout the search, potentially wasting resources on low-value candidates while overlooking promising ones. We propose EvoAlloc, a self-evolving resource-allocation agent that learns from search experience to revise its strategy for allocating computational resources across candidates. EvoAlloc periodically consolidates prior search and allocation outcomes into reusable experience, which informs subsequent strategy revisions. It further uses a counterfactual exploration mechanism to occasionally evaluate candidates denied resources by the allocator, revealing their outcomes to enrich its experience for future strategy updates. Across coding and agent-harness optimization benchmarks, EvoAlloc requires 59-82% fewer full evaluations and 61-89% fewer total LLM tokens to reach baseline-level performance. Moreover, under the same full-evaluation budget, EvoAlloc achieves 8.7-12.0% higher final performance.