MESH-Harness: Self-Improving Agent Harnesses via Bandit-Guided Compositional Evolution

πŸ“… 2026-10-04
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenge of optimizing agent frameworks under limited evaluation budgets by proposing an efficient self-improvement method that freezes model weights and optimizes only the framework. Through modular design and combinatorial evolution, it overcomes enumeration bottlenecks while synergizing module-level reuse with feedback-driven optimization via full-covariance LinUCB-guided search, mixed-init coordinate ascent, and validation-trajectory-driven code iteration. Experimental results demonstrate that the proposed approach significantly outperforms the Meta-Harness baseline across multiple tasks, reducing testing costs by 44.2% and total costs by 14.6%.
πŸ“ Abstract
An agent harness is the code that organizes context, maintains state, and coordinates tool calls for a language model. We study how to improve the harness under a limited evaluation budget while keeping model weights fixed. Our method, MESH-Harness, organizes each harness into functional modules with explicit role-specific interfaces, allowing alternative implementations of each module to be substituted and recombined. It uses shared module representations and full-covariance LinUCB to score candidate combinations based on predicted performance and exploration value. Mixed-start coordinate ascent selects complete configurations for evaluation without enumerating the combinatorial space. Validation traces then guide local code edits, and the resulting candidates are incorporated into fixed-capacity role-specific pools for subsequent recombination. On text tasks, retrieval-augmented mathematical reasoning, code generation, and interactive scientific tasks, MESH-Harness outperforms Meta-Harness by 5.70, 7.01, 2.00, and 5.00 points, respectively, under matched candidate-evaluation budgets. Iterative harness optimization improves MESH-Harness by 5.63-7.79 points over its first-round configurations. For the reported configurations, aggregate test-time cost is 44.2% lower than that of Meta-Harness, while total cost including search is 14.6% lower. These results show that combining module-level design reuse with feedback-driven compositional search can systematically improve agent harnesses while keeping overall optimization cost under control.
Problem

Research questions and friction points this paper is trying to address.

agent harness
compositional optimization
limited evaluation budget
self-improving agents
Innovation

Methods, ideas, or system contributions that make the work stand out.

Compositional Evolution
LinUCB
Agent Harness
Modular Architecture
Coordinate Ascent