LiteEvo: Automated, Cost-Efficient Harness Evolution for Generalization to Unseen Tasks

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high manual construction costs and poor generalization to unseen tasks inherent in large language model (LLM) agent frameworks by proposing a lightweight framework evolution algorithm. Starting from a neutral framework without requiring predefined benchmark names, this method employs a tool-free meta-agent, a versioned component library, and trajectory mining techniques to automatically extract reusable components from execution trajectories and dynamically compose optimized frameworks. Experimental results demonstrate that the proposed approach significantly reduces API invocation costs while substantially improving pass rates on unseen tasks and cross-task generalization capabilities. Ultimately, this work achieves efficient, automated evolution of agent frameworks, offering a scalable solution for enhancing LLM-based agents across diverse and previously unencountered scenarios.
📝 Abstract
An LLM agent is defined by two things: the weights inside its model and the harness of components assembled around it. Harnesses are still handcrafted, and HarnessX, which evolves them automatically, starts each benchmark from a handcrafted harness, reports gains on the tasks it evolved on, and budgets 100 to 175 million meta-agent tokens per benchmark. We propose LiteEvo, a lightweight harness-evolution algorithm whose tool-free meta-agents mine agent trajectories for reusable components, curate them into a versioned library, and compose each round's harness from it, starting every benchmark from the same neutral harness and never naming the benchmark. Evolving on the graded tasks of five agentic benchmarks with a frozen Qwen3.5-9B, LiteEvo lifts pass@2 by 10.5 to 67.7pp and reaches comparable or higher pass@2 than a reproduction of HarnessX (71.0 against 67.3 on average) at 13.0 lower mean API cost. Harnesses evolved on train tasks keep their gains on unseen test tasks of four benchmarks, and LiteEvo also lifts Claude Code with Sonnet 4.6 by 1.2 to 71.4pp.
Problem

Research questions and friction points this paper is trying to address.

LLM agent harness
automated evolution
generalization to unseen tasks
cost efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Harness Evolution
Meta-Agent
Trajectory Mining
Cost-Efficiency
Generalization
🔎 Similar Papers
No similar papers found.
E
Euntae Choi
Seoul National University
S
Sumin Song
Seoul National University
Sungjoo Yoo
Sungjoo Yoo
Seoul National University
memorystorage