🤖 AI Summary
This study addresses the challenge in meta-Harness methods where LLM proposers struggle to maintain evolutionary states amid expanding historical experiments. To overcome this, we propose the Causal Improvement Graph (CIG), establishing the first graph-based governance framework. This approach persistently stores evidence, hypothesis, intervention, and outcome nodes within a graph structure, externalizing improvement states through node-linking mechanisms and local operations. Consequently, subsequent optimization relies directly on structural relationships rather than tracing raw experimental histories. Empirical evaluations demonstrate that CIG consistently discovers Harness configurations superior to baselines across diverse agent tasks, while exhibiting strong robustness to the selection of both solvers and proposers.
📝 Abstract
Agentic Harness is the runtime that constructs task context and controls execution flow, thereby shaping overall agent performance. Given a fixed model and external evaluation, automated Harness optimization seeks to improve this runtime through an iterative proposal--evaluation loop to better solve target tasks. Existing meta-harness methods mainly adopt proposer-centric discovery, in which an LLM-based proposer integrates accumulated experimental findings to determine subsequent Harness revisions. This places the burden of maintaining the evolving improvement state on the proposer as history expands and its underlying experimental logic becomes harder to discern. In this paper, we introduce the Causal Improvement Graph (CIG), a graph-governed meta-harness framework that externalizes the evolving improvement state in a persistent graph, allowing prior findings to directly govern subsequent Harness optimization through local proposer operations. CIG grows and links Evidence, Hypothesis, Intervention, and Outcome nodes to represent what was observed, how it may be explained, how to test that explanation, and what the evaluation reveals. Their structural relations preserve how the improvement state changes across iterations, allowing local proposers to build directly on relations among prior findings rather than recover them from raw history. Across various agent tasks, CIG discovers stronger Harnesses than previous meta-harness baselines and remains robust to the choice of task solver and proposer. Structural ablations further support the design of an explicit improvement state with graph-governed evolution.