MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of large search spaces, high manual costs, and limited exploration in existing automated methods for designing agent system frameworks. To this end, we propose a co-evolutionary agent framework along with a novel discovery strategy. The method introduces a pioneering meta-evolutionary island orchestration mechanism that integrates multi-agent evolution, hierarchical lineage memory, mutation proxy rewriting, dynamic adaptive search, and negative evidence utilization to enable the automated and efficient generation and optimization of frameworks. Experimental results demonstrate that the proposed approach surpasses state-of-the-art frameworks across multiple benchmarks, significantly improving task resolution rates while reducing token consumption. Notably, it establishes new performance upper bounds on mathematical reasoning tasks.
📝 Abstract
Modern agentic systems combine an AI model with a harness that controls execution and environmental interactions. Harness design strongly affects long-horizon performance, yet its combinatorial search space demands substantial human effort that must be repeated as models change. Existing automated methods explore this space narrowly, optimizing only components such as prompts or skills or becoming trapped by fixed, exploitative search strategies. We introduce MILO (Meta-evolutionary Island Orchestration), a framework that co-evolves agent harnesses and the strategy used to discover them. MILO combines: (i) hierarchical lineage memory over island-based trees, using rejected mutations as negative evidence; (ii) per-island mutator agents that rewrite complete harnesses using global search history and parent-specific feedback; and (iii) an orchestrator that adapts search through lineage grafting and speciation, mutator reassignment and curriculum revision. Across Terminal-Bench 2.1, PaperBench, and DeepSWE, MILO-discovered harnesses outperform eight state-of-the-art harnesses and six search methods using frontier (Opus 4.8) and open-weight (gpt-oss-120b) models. With Opus 4.8, MILO improves resolution over its initial harness by $+12.0\%$, $+28.3\%$, and $+10.3\%$, respectively, compared with best prior-search gains of $+4.5\%$, $+18.3\%$, and $0\%$. On Terminal-Bench 2.1, it achieves $86.1 \pm 2.0\%$, exceeding the official leaderboard's top entry ($83.8 \pm 2.3\%$) while using 26\% fewer tokens than its initial harness. On EinsteinArena open problems, MILO improves best-known upper bounds for Erdős minimum-overlap ($0.3808586 \to 0.3808568$) and the first and third autocorrelation inequalities ($1.50274365 \to 1.50274360$; $1.45081 \to 1.44889$).
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Agent Evolution
Harness Discovery
Meta-evolutionary Island Orchestration
Hierarchical Lineage Memory
Co-evolution
🔎 Similar Papers
No similar papers found.