test harness development

Designs, implements, and maintains evaluation harnesses—end-to-end test frameworks and tooling that generate inputs/queries, control task and environment variables, run agents under specified conditions, collect metrics, and automate benchmarking. Uses harness-aware and environment-shift benchmarking to build multi-factor, out-of-distribution, and query-based evaluations that measure performance across harness variants, reveal interaction effects between system design and the evaluation setup, and identify failure modes for analysis and reporting.

testharnessdevelopment

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
3.56
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$202K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Current agent evaluation practices often reduce failures to system-level outcomes, making it difficult to pinpoint root causes or guide effective remediation. This work proposes an interaction-centric failure taxonomy and introduces, for the first time, a cross-architectural and generalizable framework for failure localization. The framework maps 41 distinct failure modes onto interaction edges between components—such as models, toolchains, and environments—and explicitly delineates responsibility boundaries among them. By integrating component interaction graph attribution, multi-source trajectory analysis, and an independent reasoning agent-based evaluator, the approach enables reproducible validation. Experiments across four state-of-the-art models demonstrate that the strongest evaluator achieves a Cohen’s κ of 0.76 with human annotations, confirming the taxonomy’s generalizability and consensus alignment.

agent failuresfailure localizationinteraction-centric taxonomy

Must-Read Papers

Most classic and influential ideas
View more

Existing benchmarks struggle to evaluate the impact of harness components on the performance of large language model (LLM) agent systems, often overlooking execution details or fixing harness configurations. This work proposes Harness-Bench, the first framework to systematically decouple and quantify how harness design influences agent workflows. By enforcing a unified task environment, computational budget, and evaluation protocol—and leveraging sandboxed offline tasks, realistic usage patterns, human auditing, and full execution trace logging—it enables controlled experiments across diverse model–harness combinations. Analysis of 5,194 execution traces across 106 tasks reveals significant differences in completion rates, process quality, efficiency, and failure modes, exposing alignment failures stemming from misalignment between reasoning and execution. The study argues that agent performance should be reported based on joint model–harness configurations rather than base models alone.

agent workflowsbenchmarkingexecution-layer variation

This work addresses the limitations of existing automatic harness evolution methods, which are prone to evaluation overfitting due to overlap between search and test benchmarks and lack fair comparison against test-time search baselines under identical feedback and reasoning budgets. To remedy this, the study introduces a rigorously controlled evaluation protocol on Terminal-Bench 2.1, employing GPT-5.4 and Claude Opus 4.6. Through budget-matched test-time scaling baselines, controlled ablation studies, and held-out task generalization assessments, the empirical analysis reveals that performance gains attributed to automatic harness evolution primarily stem from additional search rather than algorithmic improvements. These gains prove non-persistent and exhibit limited generalization, casting doubt on the efficacy of current approaches.

evaluation protocolgeneralizationharness evolution

This work addresses the current lack of a standardized protocol for evaluating large language models’ (LLMs’) ability to optimize external components of intelligent agents—such as prompts, tools, and control flows. It introduces the first auditable, resource-constrained, and evaluation-isolated harness optimization benchmarking framework, which enforces assessment boundaries via trusted execution environments and incorporates standardized scoring, version tracking, and fixed-budget controls to enable systematic multi-model, multi-task, and multi-seed experimentation. Across 111 experimental runs, the study reveals that the optimizer model itself is more discriminative than the initial harness, that native harnesses exhibit no consistent advantage, and that optimization gains are highly dependent on both task and initial configuration. This work establishes harness optimization as a measurable and discriminative capability in AI systems.

agentic systemsautomated optimizationbenchmarking

This study addresses the significant yet underexplored impact of framework design on LLM agent performance in software engineering. We present the first systematic quantification of framework effects, empirically analyzing the interaction mechanisms among core components—including tool registration, context compression, and sub-agents—using the SWE-bench benchmark with Qwen and DeepSeek models. To facilitate this analysis, we introduce NanoHarness, a lightweight, modular evaluation framework. Our findings establish framework design as a primary determinant of agent performance and reveal diminishing marginal returns when applying complex frameworks to highly capable models. Notably, NanoHarness successfully replicates the performance gains achieved by most production-grade frameworks, yielding a substantial 7.37% improvement in agent effectiveness.

Agent HarnessEmpirical StudyHarness Design

Latest Papers

What's happening recently
View more

This study addresses the limited adaptability of fixed agent frameworks and the high cost and risks associated with dynamic code generation by proposing STITCH. This framework decouples framework construction from code generation by mining reusable primitives from failure trajectories, which are dynamically selected and compiled into task-specific frameworks at test time based on task information, thereby eliminating the need for online debugging. Experimental results demonstrate that STITCH improves success rates by 12% over Codex CLI while incurring only a 2.7% composition overhead. Furthermore, it achieves a 638-fold efficiency gain compared to generating solutions from scratch, significantly enhancing both system robustness and execution efficiency.

Agent HarnessLarge Language ModelsReusable Primitives

This study addresses the limitation of existing globally unified testing frameworks, which achieve average optimality yet remain suboptimal for individual instances and struggle to adapt to specific task cases. To overcome this, we propose the first adaptive framework that recycles information from global optimization experiences. Methodologically, our approach repurposes global optimization artifacts to generate structured manuals and trains an editor to craft instance-aware patches for each case, thereby enabling dynamic framework generation and optimization. We evaluate the proposed method across seven benchmarks encompassing interactive agents, software engineering, and long-horizon terminal tasks. Experimental results demonstrate that our approach consistently outperforms existing baselines, establishing a robust solution for instance-level adaptation in complex testing environments.

AI agentsharness optimizationinstance-adaptive

This study addresses the insufficient test coverage of agent frameworks and the low reliability of LLM-dependent code by presenting the first empirical investigation into testing adequacy for such frameworks, alongside a novel technique termed HarnessTester. This approach generates faithful test suites based on explicit agent-framework contracts and integrates static analysis with dynamic execution to substantially enhance both the depth and breadth of testing for LLM-dependent code. Experimental results demonstrate that HarnessTester effectively improves line and branch coverage as well as mutation scores. Furthermore, it identifies 122 real-world bugs, including 88 previously unknown defects, 69 of which have already been confirmed by developers.

Agent HarnessLLM-based AgentsSoftware Reliability

This study addresses the reliance of multi-agent system architectures on costly execution or manual design by proposing the SHIFT framework. SHIFT decouples execution from the search loop, employing a local LLM as an architect to learn construction policies and integrating a value function for utility prediction. Through Monte Carlo Tree Search, it dynamically generates optimal architectures per query, enabling joint optimization of structure, instructions, and tools. Experimental results demonstrate that SHIFT achieves an average accuracy of 80% across six benchmarks, outperforming the strongest baseline by 7.2 percentage points while reducing execution tokens by 32%, thereby significantly balancing performance and computational cost.

Agent HarnessDynamic SearchExecution Cost

This study addresses the limited generalizability of existing self-evolution frameworks, which rely on independent proposers and are confined to single benchmarks. We propose a recursive self-improvement framework enabling a single frozen model to simultaneously serve as both solver and optimizer within a unified harness, achieving multi-task generalization by directly editing its own code. This approach formulates evolution as a two-stage process comprising multi-task pretraining and continual learning, integrating techniques such as LLM-agent self-optimization, multi-benchmark co-evolution, and history compression to eliminate reliance on external human-designed harnesses. Experiments demonstrate that the evolved seed harness yields average improvements of 4.48 points in-distribution and 12.64 points out-of-distribution, surpassing Codex. Furthermore, continual evolution establishes new state-of-the-art records on specific benchmarks.

multi-task generalizationout-of-distribution evaluationrecursive self-improvement

Hot Scholars

DL

David Lo

Professor of Computer Science, Singapore Management University
AI4SESoftware AnalyticsSE4AISoftware Maintenance
XW

Xingyao Wang

All Hands AI, University of Illinois Urbana-Champaign
ZZ

Zibin Zheng

IEEE Fellow, Highly Cited Researcher, Sun Yat-sen University, China
BlockchainSmart ContractServices ComputingSoftware Reliability
RJ

Reyhaneh Jabbarvand

Siebel School of Computing and Data Science, University of Illinois at Urbana-Champaign
Neuro-symbolic program analysisCode LLMs (evaluationinterpretabilityand benchmarking)