Self-Evolving Harness on Multiple Tasks with the Agent as Its Own Optimizer

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited generalizability of existing self-evolution frameworks, which rely on independent proposers and are confined to single benchmarks. We propose a recursive self-improvement framework enabling a single frozen model to simultaneously serve as both solver and optimizer within a unified harness, achieving multi-task generalization by directly editing its own code. This approach formulates evolution as a two-stage process comprising multi-task pretraining and continual learning, integrating techniques such as LLM-agent self-optimization, multi-benchmark co-evolution, and history compression to eliminate reliance on external human-designed harnesses. Experiments demonstrate that the evolved seed harness yields average improvements of 4.48 points in-distribution and 12.64 points out-of-distribution, surpassing Codex. Furthermore, continual evolution establishes new state-of-the-art records on specific benchmarks.
📝 Abstract
A harness is the code around a language-model agent that organizes prompts, calls tools, manages context, and controls execution. As models grow stronger, recent work has begun to let agents improve their own harnesses, a line of work known as self-evolving harnesses. In most existing methods, a separate proposer running on a human-designed harness modifies the solver's harness, and a separate harness is evolved for each benchmark. Real-world tasks come from many domains, so both the evolution and the evaluation of a harness should cover a diverse range of tasks. We propose a framework close to recursive self-improvement: the same frozen model, on the same version of the harness, first solves tasks as the solver and then, as the proposer, reads the complete run records and directly edits the harness that runs it. Each evolution batch draws tasks from five benchmarks in different domains. To measure generalization, training and held-out tasks are strictly separated, and we additionally evaluate on five out-of-distribution benchmarks never used during evolution. We frame the evolution process as deep-learning training with two stages, multi-task pretraining and continual training. Starting from a 49-line seed harness, the harness obtained at the end of the first stage improves the average score by 4.48 points on the in-distribution benchmarks and by 12.64 points on the out-of-distribution benchmarks, surpassing Codex on the former and matching it on the latter. In the second stage, continued evolution on Claw-Eval, one of the out-of-distribution benchmarks, further raises the score on that benchmark from 66.17 to 68.06, exceeding Codex. We also provide an in-depth analysis of the mechanisms that emerged during evolution, including output truncation, history compaction, and independent review.
Problem

Research questions and friction points this paper is trying to address.

self-evolving harness
multi-task generalization
recursive self-improvement
out-of-distribution evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-evolving harness
Recursive self-improvement
Multi-task pretraining
Out-of-distribution generalization
Continual training
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Q
Qiankai Xu
Nanjing University