Harness Engineering for Software Engineering via Modular Executable Dev-Primitives

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of context explosion and semantic drift caused by state reconstruction when large language model (LLM) agents perform long-horizon software engineering tasks. To this end, we propose Dev-Primitives, a novel abstraction that transforms static code components into active primitives with autonomous reasoning capabilities, along with the HERMES framework. This approach introduces a dependency-aware dynamic activation mechanism and fault diagnosis mapping, rendering software artifacts agent-native interfaces. Furthermore, it constructs a comprehensive execution pipeline by integrating terminal access, dependency analysis, and local self-modification algorithms. Experimental evaluations across four benchmarks demonstrate that HERMES achieves an average performance improvement of 12.4% while significantly reducing inference costs, thereby validating its critical value for advancing LLM-based software engineering agents.
📝 Abstract
Large language models (LLMs) equipped with terminal access have demonstrated strong capabilities in automating software engineering tasks. However, existing agents remain brittle on long-horizon workflows, where they must repeatedly reconstruct program state scattered across source files, configurations, tests, dependencies, and runtime behavior, leading to increasingly long interaction histories, context explosion, and semantic drift. Large repositories further complicate the identification of task-relevant components. To address these challenges, we introduce \textbf{Dev-Primitives} (\emph{Development Primitives}), a modular and executable abstraction that transforms repository components from passive software artifacts into active participants in software engineering. Each Dev-Primitive pairs a repository artifact with a resident LLM, which gives the artifact an agent-native interface grounded in its own implementation and dependencies, enabling natural-language reasoning, inter-component communication, and localized self-modification. Building on Dev-Primitives, we propose \textbf{HERMES}, a Harness Engineering framework for software engineeRing via Modular Executable Dev-PrimitiveS, which instantiates these primitives at repository scale through a dependency-aware dynamic activation mechanism and a bug diagnosis mechanism that maps execution evidence back to the components that must be revised. Extensive experiments on four software engineering benchmarks demonstrate that HERMES outperforms matched baseline harnesses by 12.4\% on average. Moreover, when paired with strong activation and diagnosis models, HERMES, even with Qwen3-8B Dev-Primitives, remains within 4.5\% of the homogeneous GPT-5.6 Sol configuration across all four benchmarks, while reducing inference cost by 26.2\% on Terminal-Bench 4.0, highlighting the importance of harness design in software engineering agents.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Software Engineering Agents
Context Explosion
Semantic Drift
Long-horizon Workflows
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dev-Primitives
Modular Executable Abstraction
Harness Engineering
Dependency-aware Dynamic Activation
Software Engineering Agents