๐ค AI Summary
This study addresses the efficiency bottlenecks of LLM agents in long-horizon automated research, which arise from history replay, redundant work, and premature stagnation. To overcome these challenges, this work proposes a stateless language agent framework that introduces a novel โstateful search with stateless agentsโ paradigm. Specifically, it decouples persistent research states from dialogue histories by delegating state management to external tools that dynamically reconstruct context for each invocation. Furthermore, an Advisor-Worker architecture is designed to optimize resource scheduling through parallel execution and centralized evidence summarization. Experimental results demonstrate that the proposed approach achieves state-of-the-art performance on tasks such as software engineering while reducing token consumption by over 84%. Additionally, this work reveals inherent limitations associated with short evaluation horizons.
๐ Abstract
Automated research systems increasingly run LLM agents over long horizons, but more inference does not by itself produce more progress: agents replay growing histories, duplicate one another's work, or stop experimenting while token consumption continues. Yet most evaluations use short budgets or benchmarks that saturate early, leaving these failure modes untested. We trace these failures to two choices: where research state lives and who decides what to try next. We introduce Stateless Language Agents (SLAs), built on the principle of stateful search with stateless agents: no agent carries its conversation across invocations; instead, the harness owns the research state (candidate solutions and measured outcomes) and reconstructs a fresh and role-specific context for every invocation. What each agent sees becomes an explicit design choice rather than a history that grows with the run. We implement this principle in the SLA framework, where a stateless Advisor reads harness-summarized evidence across search directions and assigns concrete experiments to parallel Workers. We evaluate SLA against three recent frameworks on software engineering, kernel optimization, and algorithm design at budgets of up to one billion tokens. SLA achieves the best final result on every task and reaches the strongest kernel baseline's final performance with over 84% fewer tokens. Ablations from shared checkpoints show that focused contexts and explicit assignments each contribute to SLA's progress, with effects that can compound over full runs, while the Advisor consumes less than 0.6% of tokens. These results argue for SLAs, which keep durable research state out of agent conversations, and show that short evaluation horizons can misjudge research systems and their components.