Asynchronous Is Nearly Free for Evolution Strategies on Long-Horizon Agentic Tasks

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the waiting bottlenecks of synchronous evolution strategies (ES) and the staleness issues of asynchronous approaches in training long-horizon LLM agents. We propose a bounded-staleness asynchronous ES method based on Qwen2.5-7B-Instruct, which restricts policy update delays to enable efficient post-training optimization within an asynchronous reinforcement learning framework. This work provides the first empirical validation of both synchronous and asynchronous ES on multi-turn terminal coding tasks, revealing that ES exhibits strong robustness to moderate policy staleness. Experimental results demonstrate that asynchronous ES achieves a 25.9% success rate, marginally outperforming the 25.4% attained by its synchronous counterpart. Furthermore, introducing delay incurs only negligible performance degradation, confirming that asynchronization is virtually lossless while significantly reducing training latency and improving trajectory data utilization.
📝 Abstract
LLM-based long-horizon agentic post-training is often bottlenecked by rollout generation: trajectories span many interaction turns, completion times vary substantially, and synchronous update barriers leave faster workers waiting for stragglers. Asynchronous reinforcement learning which has been adopted in LLM post-training addresses this inefficiency by consuming trajectories as they arrive, but introduces policy lag and off-policy optimization. Evolution strategies (ES) offer a backpropagation-free alternative for LLM post-training, yet it relies on a larger number of rollouts and existing practices have remained largely synchronous. In this short-form paper, we introduce bounded-staleness asynchronous ES and demonstrate it on Endless Terminals benchmark using Qwen2.5-7B-Instruct. Across three evaluation seeds, natural Async-1 matches synchronous ES, achieving 25.9\% versus 25.4\% held-out success. Controlled schedules that delay 10\% of each update cohort by four or eight policy updates reduce success by only 1.6 and 3.1 percentage points, respectively, without explicit off-policy correction. GRPO performs better overall, reaching 29.0\% held-out success, but importantly our results show that ES tolerates moderate policy staleness with limited degradation, opening possibilities for future improvement of ES-based post-training with asynchronous algorithms. To the best of our knowledge, we are the first to demonstrate the effectiveness of sync and async ES on a multi-turn terminal style agentic coding task.
Problem

Research questions and friction points this paper is trying to address.

Evolution Strategies
Long-horizon Agentic Tasks
Asynchronous Training
LLM Post-training
Policy Staleness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Asynchronous Evolution Strategies
Bounded Staleness
LLM Post-training
Long-Horizon Agentic Tasks
Off-policy Tolerance
🔎 Similar Papers
No similar papers found.