Beyond the Model: Demystifying Harness Effects in Software Engineering Agents

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the significant yet underexplored impact of framework design on LLM agent performance in software engineering. We present the first systematic quantification of framework effects, empirically analyzing the interaction mechanisms among core components—including tool registration, context compression, and sub-agents—using the SWE-bench benchmark with Qwen and DeepSeek models. To facilitate this analysis, we introduce NanoHarness, a lightweight, modular evaluation framework. Our findings establish framework design as a primary determinant of agent performance and reveal diminishing marginal returns when applying complex frameworks to highly capable models. Notably, NanoHarness successfully replicates the performance gains achieved by most production-grade frameworks, yielding a substantial 7.37% improvement in agent effectiveness.
📝 Abstract
Large Language Model (LLM)-based agents are increasingly used for software engineering tasks, yet their performance is not determined by the base model alone. The agent harness substantially shapes how SE agents interact with repositories, execute actions, and validate solutions. However, the role of harness design remains insufficiently understood, especially across different models, tasks, and harness components. In this paper, we present a systematic empirical study of harness effects in SE agents. We first evaluate two representative harnesses, mini-SWE-agent and OpenCode, with ten models from two prominent open-weight model families, Qwen and DeepSeek, on three benchmarks: SWE-bench Pro, ProgramBench, and GitTaskBench. We then construct NanoHarness, a lightweight modular harness built on top of mini-SWE-agent, and use it to analyze five representative harness components: tool registry, context compression, explicit planning, subagents, and lazy skills. Experimental results show that harness effectiveness depends jointly on model capability and task type. Complex harnesses provide diminishing marginal gains on SWE-style issue repair as model capability improves, but can benefit stronger models on more complex and open-ended repository-level tasks. Component-level analysis on ProgramBench further shows that structured tool use and task-specific subagents provide the most stable improvements, while context compression and general subagents can hurt repository-generation performance. When combined, NanoHarness improves over mini-SWE-agent by 7.37 and 6.21 percentage points on Qwen3.7-Max and DeepSeek-V4-Pro, respectively, recovering most of the gains of product-level harnesses. These findings highlight harness design as a first-class factor in SE-agent performance and provide insights for building more effective and efficient coding agents.
Problem

Research questions and friction points this paper is trying to address.

Software Engineering Agents
Agent Harness
Large Language Models
Harness Design
Empirical Study
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agent Harness
NanoHarness
Modular Framework
Software Engineering Agents
Component-level Analysis
🔎 Similar Papers
No similar papers found.