Certified Long-Horizon Code Agent Evolution via Validation-Gated Skill Optimization

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the problem of performance regression and system collapse caused by context self-evolution in long-horizon code agents operating on repository-level tasks. To mitigate this, we propose VALVE, a validation-gated framework that formalizes weight-free context self-evolution by integrating textual skill refinement, statistical hypothesis testing, and LLM-based agent techniques. The framework provides finite convergence guarantees along with rigorous theoretical bounds on future task gains and regressions. Evaluated on over one thousand SWE-bench tasks, VALVE achieves an average improvement of 14.9 points, reduces the regression rate by 75%, and compresses the skill library by 11×, thereby enabling stable and reliable long-term performance evolution.
📝 Abstract
Long horizon agent self-evolution without model weight updates is essential for enabling deployed agents to accumulate reusable skills and improve over time. Prior self-evolution work has focused primarily on short-horizon tasks, while repository-level software engineering remains unexplored despite being an ideal testbed for long-horizon adaptation. In this setting, agents are required to solve streams of sequential tasks, navigate complex dependencies with evolving repositories and persistently store and reuse experience. Text-based skill optimization offers an efficient, non-parametric approach for such adaptation. However, existing methods often suffer from unstable updates, performance drawdown, and agent collapse over extended deployments. In this paper, we formalize the concept of in-context self-evolution and introduce VALVE, a validated-gated framework for long-horizon skill optimization. We establish finite convergence, provide theoretical guarantees for future-task gain and drawdown, and derive the validation and evaluation holdout sizes required for a prescribed tolerance, with leading-order scaling Empirically, our pipeline, VALVE achieves stable self-improvement over evolution horizon spanning more than 1,000 SWE tasks, with average final and peak gains of $14.9$ and $16.5$ points across three frontier models (GPT-5.5, Claude-4.6 and MiniMax-M2.7). The validation gate reduces average drawdown by 75% and produces an 11x more compact skill bank than ungated evolution. We further present extensive ablations identifying the design choices most critical to long-horizon skill evolution.
Problem

Research questions and friction points this paper is trying to address.

long-horizon self-evolution
code agent
skill optimization
in-context learning
performance drawdown
Innovation

Methods, ideas, or system contributions that make the work stand out.

In-context self-evolution
Validation-gated skill optimization
Long-horizon code agent
Non-parametric adaptation
Finite convergence guarantees