🤖 AI Summary
This work addresses the critical yet often overlooked dependence of agent performance on harness engineering, where control logic is typically entangled within code, hindering transferability, reuse, and systematic study. To overcome this limitation, we propose— for the first time—externalizing the high-level control logic of harnesses into editable, executable natural language specifications, supported by a unified Intelligent Harness Runtime (IHR) architecture. The IHR introduces explicit contracts, persistent artifacts, and lightweight adapters to enable modular harness design and cross-task transfer. We validate our approach on programming and computer-use benchmarks, demonstrating its effectiveness through comprehensive experiments. Ablation studies and successful transfers from code-based to text-based harnesses further illustrate the framework’s flexibility and feasibility.
📝 Abstract
Agent performance increasingly depends on \emph{harness engineering}, yet harness design is usually buried in controller code and runtime-specific conventions, making it hard to transfer, compare, and study as a scientific object. We ask whether the high-level control logic of an agent harness can instead be externalized as a portable executable artifact. We introduce \textbf{Natural-Language Agent Harnesses} (NLAHs), which express harness behavior in editable natural language, and \textbf{Intelligent Harness Runtime} (IHR), a shared runtime that executes these harnesses through explicit contracts, durable artifacts, and lightweight adapters. Across coding and computer-use benchmarks, we conduct controlled evaluations of operational viability, module ablation, and code-to-text harness migration.