Self-Evolving Algorithm-Design Agents: Escaping In-Context Evolutionary Stagnation via Population-Curated Policy Optimization

๐Ÿ“… 2026-09-29
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the challenges of evolutionary stagnation caused by domain knowledge gaps and parameter adaptation under high-quality data scarcity in algorithm design agents. To this end, we propose a sample-efficient parameter self-evolution framework that leverages self-generated algorithms for exploratory learning. By integrating an enhanced Chain Proposition and Population Curation Optimization (PCPO) strategy with a hybrid update mechanism, the framework enables efficient knowledge internalization and reuse. Experimental results demonstrate that an 8B-parameter model employing our approach surpasses state-of-the-art methods on chip placement tasks while achieving performance comparable to leading closed-source models. Furthermore, it yields an average 8.27ร— speedup in GPU kernel design and significantly reduces inference costs.
๐Ÿ“ Abstract
Large language models are increasingly participating in complex real-world tasks in the form of algorithm-design agents, designing and refining algorithms. Many successful algorithm-design agents adopt pure in-context evolutionary frameworks, but they may quickly plateau in domains that require specialized knowledge. Parametric adaptation offers a way to internalize specialized knowledge, but conventional training requires abundant domain-specific corpora while high-quality algorithms are scarce in complex algorithm-design scenarios. In this paper, we propose sample-efficient parametric self-evolution where agents can explore and learn from self-generated algorithms. First, we characterize in-context evolutionary stagnation and analytically propose the Improvement Chain proposition, showing how learning successive self-generated algorithms can locally increase the likelihood of neighboring algorithms. Motivated by this local-transfer perspective, we further propose Population-Curated Policy Optimization (PCPO) to utilize a global population and a hybrid policy update scheme for retaining and reusing high-quality, diverse self-generated algorithms, shifting the policy towards stronger algorithms. In the task of learning rate schedule design for global placement in electronic design automation, trained only on 4 chip cases, PCPO outperforms the state-of-the-art in-context evolutionary methods (e.g., OpenEvolve and ShinkaEvolve) on average across 16 chip cases. With an 8B-size base model, PCPO achieves competitive performance compared to frontier closed-source models such as GPT-5.5. PCPO also reduces inference-time token cost by internalizing grounded domain knowledge and prompt distillation. Moreover, PCPO achieves significant speedups on four GPU kernel designs, with an average of 8.27$\times$ speedup against the PyTorch Eager baseline.
Problem

Research questions and friction points this paper is trying to address.

algorithm-design agents
in-context evolutionary stagnation
parametric adaptation
sample-efficient self-evolution
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Evolving Agents
Policy Optimization
In-Context Evolution
Algorithm Design
Sample Efficiency
C
Chen Lu
State Key Laboratory of Novel Software Technology, Nanjing University, China; School of Artificial Intelligence, Nanjing University, China
Ke Xue
Ke Xue
Nanjing University
Black-Box OptimizationMachine Learning
S
Siyuan Xu
Huawei Noahโ€™s Ark Lab, China
M
Mingxuan Yuan
Huawei Noahโ€™s Ark Lab, China
Chao Qian
Chao Qian
Nanjing University
Artificial intelligenceevolutionary algorithmsmachine learning