ForkPilot: Self-Evolving Policy for Retrospective Search in Long-Horizon Agents

๐Ÿ“… 2026-10-04
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the resource misallocation problem in long-horizon agent backtracking search caused by the complexity of attribution and adaptation. To this end, we propose ForkPilot, a self-evolving strategy framework that introduces a dynamic search value model. This framework establishes a two-stage "offline learningโ€“online updating" mechanism integrating reinforcement learning, automated trajectory contrastive analysis, and large language model reasoning to achieve a paradigm shift from static estimation to dynamic adaptation. Experimental results demonstrate that ForkPilot attains state-of-the-art performance across multiple benchmarks while significantly reducing token consumption by 59.2%, thereby substantially improving search efficiency.
๐Ÿ“ Abstract
Interactive language-model agents increasingly solve complex tasks through long-horizon, multi-call reasoning, where errors in beliefs or actions can compound across tool interactions. Retrospective search can recover from such failures but is prone to misallocation. Delayed outcomes obscure the contribution of intermediate search decisions, leading to Attribution Complexity, while evolving execution evidence leads to Adaptation Complexity, where previously learned estimates become stale. To address these challenges, we first introduce Search Value Dynamics (SVD), which characterizes the evolving trade-off between the gain and cost of retrospective search. Building on SVD, we propose ForkPilot, a self-evolving two-stage policy-learning framework. In the first stage, ForkPilot learns a search-value policy offline from completed trajectories through automatically constructed outcome comparisons. In the second stage, it makes search decisions based on current observations and then self-evolves by incorporating newly completed trajectories into subsequent policy updates. We evaluate ForkPilot across 6 diverse benchmarks and 7 widely used LLM backbone families, including four open-source families, GPT-5.6 Sol, and Opus 4.8 in a production agentic system, against 9 competitive baselines, including a real-world harness deployment used by hundreds of thousands of paid users. ForkPilot achieves comparable state-of-the-art performance while reducing token usage by up to 59.2%, demonstrating its efficacy.
Problem

Research questions and friction points this paper is trying to address.

long-horizon agents
retrospective search
attribution complexity
adaptation complexity
error compounding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Evolving Policy
Retrospective Search
Search Value Dynamics
Long-Horizon Agents
Policy Learning
๐Ÿ”Ž Similar Papers