π€ AI Summary
This study addresses the limitations of existing self-improvement approaches for large language model (LLM) agents, which typically rely on downstream evaluation and are thus constrained by high computational costs and task-specific dependencies. To overcome these challenges, this work proposes SelfSearch, a novel framework that introduces the first reward-free search mechanism grounded in historical experience. By integrating the coding capabilities of LLMs with a self-modifying agent architecture, SelfSearch leverages experience replay to guide agents in autonomously optimizing their instructions and tools without relying on external reward signals. This approach significantly reduces computational overhead while enhancing execution efficiency. Empirical evaluations demonstrate that SelfSearch achieves up to an 11.2% improvement in success rate on the Terminal-Bench benchmark and reduces operational costs by 38.5% on SWE-bench, highlighting its effectiveness as a scalable solution for autonomous agent improvement.
π Abstract
Advances in the coding capabilities of LLM agents allow them to inspect and modify their own instructions, tools, and execution procedures. Existing approaches use this ability to search for improved agents through repeated downstream evaluation, which incurs substantial costs and ties the search to the evaluated tasks. We introduce \textbf{SelfSearch}, a reward-free search procedure in which agents modify themselves using records of previous self-improvement episodes. These records capture the reasoning, tool actions, and outcomes of earlier modification attempts, providing concrete experience for improving both task solving and self-modification. Without downstream reward signals during search, SelfSearch improves population-mean success over the initial agent in all six model--benchmark settings, with individual agents gaining up to 11.2 percentage points on Terminal-Bench 2.1. On SWE-bench Multilingual, an agent improves success by \textbf{5.0} percentage points while reducing execution cost by \textbf{38.5}\% on tasks solved by both the initial and evolved agents. SelfSearch achieves competitive task success with evaluation-guided search baselines at lower search cost. With only \textbf{\$4.03} in search cost, it produces a harness that solves \textbf{82.0}\% of Terminal-Bench 2.1 tasks with DeepSeek V4 Flash under the settings of a public nine-harness comparison, matching the top-scoring harness, Codex. These results suggest that experience gained through self-modification can improve agents' downstream capabilities and efficiency.