Beyond Skill Evolution: Self-Evolving Context Management Policies for Long-Horizon Agent Harnesses

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the performance degradation of LLM-based agents in long-horizon tasks caused by contextual redundancy and outdated information. To this end, it proposes ContextEvo, a novel framework that formulates context management as an evolvable strategy for the first time. Built upon the Pi-agent architecture, ContextEvo optimizes adaptive information visibility by learning from execution trajectories, integrating key decision point reconstruction with directed policy updates to overcome the limitations of conventional skill evolution during extended interactions. Experimental results demonstrate that the proposed approach significantly outperforms mainstream agents, including Codex, across three benchmarks, effectively alleviating long-horizon information pressure.
📝 Abstract
Harness evolution improves LLM agents by learning from execution trajectories, but existing experience- and skill-based methods are less effective on long-horizon tasks. As interactions grow, useful evidence can be buried by redundant or outdated context, making context management itself a key bottleneck. We introduce ContextEvo, a framework that learns a context policy from long-horizon trajectories. ContextEvo reconstructs the model-visible context at key decision points, identifies context-related failures, and applies targeted policy updates. Starting from the open-source Pi-agent harness, ContextEvo improves performance across three long-horizon task benchmarks, achieving results comparable to or better than several prominent agent harnesses, including Codex, OpenCode, and OpenClaw. Additional analyses show that fixed or locally evolved context strategies can fall short under long-horizon information pressure, while our methods adapt to the information demands of each environment.
Problem

Research questions and friction points this paper is trying to address.

long-horizon tasks
context management
LLM agents
harness evolution
information pressure
Innovation

Methods, ideas, or system contributions that make the work stand out.

Context Management
Self-Evolving Policy
Long-Horizon Tasks
LLM Agents
Trajectory Learning