Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements

📅 2026-08-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出Agentic ESOpt,用进化策略方法解决长时序LLM代理微调问题,克服了RL在处理大规模模型和长时序任务上的局限性。
📝 Abstract
Reinforcement Learning (RL) has been promising in single-turn LLM fine-tuning. However, long-horizon agentic reasoning introduces increasingly branching interactions and sparse rewards, exposing several limitations of RL: its heavyweight backpropagation-based training stack makes it impractical to fine-tune larger LLMs, and longer-horizon trajectories make credit assignment in RL substantially harder. This paper argues that evolution strategies (ES) can be a better choice for fine-tuning long-horizon LLM agents. Compared with agentic RL, ES offers three key advantages: 1) Model Scalability: ES enables full-parameter optimization with only minimal, inference-level GPU memory, making it possible to fine-tune large LLMs. 2) Flexibility: its lightweight, black-box feedback interface makes ES fine-tuning easy to compose with prompt-space evolution (e.g., skill optimization & test-time compute); and 3) Long-Horizon Scalability: ES performs trajectory-level parameter attribution without decomposing rewards across horizons, yielding better scalability than Agentic RL as the horizon length grows. Based on this insight, we propose Agentic ESOpt, a full-parameter agentic fine-tuning framework tailored to flexible parameter--context co-evolution. At each step, Agentic ESOpt samples perturbations around the current LLM parameters, evaluates the resulting agents with rewards, and applies an online reward-weighted update. To improve the exploration--adaptation trade-off, Agentic ESOpt further introduces a cosine decay schedule of the perturbation scale $σ$. On WebArena-Lite, full-parameter optimization of Qwen-3.5-27B improves the No Skill baseline by 6.69%. In test-time automatic heuristic design, Agentic ESOpt performs online prompt--parameter co-evolution, improving its matched baseline in 28 of 36 settings.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Long-Horizon LLM Agents
Evolution Strategies
Fine-Tuning
Scalability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Evolution Strategies
Long-Horizon LLM Agents
Full-Parameter Optimization
Cosine Decay Schedule
🔎 Similar Papers
No similar papers found.