Score
Designs and builds the runtime harness and tooling scaffold that connects an autonomous agent to external tools, I/O channels, and control surfaces. Work includes selecting which tools to expose, authoring tool descriptions and metadata, specifying auxiliary per-step observations, and configuring control interfaces and information granularity.
This study addresses the question of whether performance bottlenecks of large language model (LLM) agents in long-horizon tasks stem from the base model, the execution framework, or their coupling. To this end, the authors propose a “model–framework” co-analysis perspective, decomposing the execution framework into six runtime responsibilities: observation, context, control, action, state, and verification. They further integrate this decomposition with the four paradigms of agent engineering—prompt engineering, workflow design, context engineering, framework engineering, and co-training—to systematically categorize existing approaches. The resulting analytical framework elucidates how runtime design choices influence task success rate, efficiency, and reliability, while highlighting critical challenges including value-aware evaluation, safety, framework generalization, and co-evolution of models and frameworks.
This work addresses the limitations of large language model agents in reliably executing complex tasks due to inherent cognitive burdens that cannot be overcome by parametric capabilities alone. The authors propose a system-level framework centered on externalization, treating memory, skills, and protocols as three interdependent cognitive artifacts coordinated through a unified execution architecture. Grounded in cognitive artifact theory, the framework integrates modular external components with runtime coordination mechanisms, offering a cohesive analytical lens for agent infrastructure. The study demonstrates that performance gains increasingly stem from external cognitive infrastructure rather than model scaling alone, highlighting promising directions such as self-evolving architectures and shared infrastructures, while also identifying associated challenges in evaluation and governance.
This work addresses the limited reliability of autonomous software engineering agents in real-world settings, often attributed to inherent model limitations. It proposes the AI Harness Engineering framework, which conceptualizes software engineering capability as a synergistic system comprising the model, a control layer, and the environment. The framework formally defines eleven core responsibilities of the AI control layer for the first time and introduces a four-tiered (H0–H3) runtime support architecture alongside a trajectory-based, auditable evaluation protocol. By integrating key techniques—such as task specification, context selection, tool access, and project memory—the framework generates structured evidence bundles in controlled tasks, enabling higher-level control layers to produce reproducible logs, failure attribution reports, determinism checks, and verification artifacts. This significantly enhances the verifiability and maintainability of code changes.
This study addresses the conceptual ambiguity surrounding the term “agent harness,” which currently lacks a unified and operational definition, leading to inconsistent usage in both research and practice. Through conceptual analysis, terminological lineage reconstruction, engineering documentation review, and empirical case validation, this work proposes the first necessary and sufficient constitutive definition of an agent harness. Building on this definition, the authors establish inclusion and exclusion criteria that enable consistent classification across six real-world systems and boundary cases. The resulting framework effectively clarifies the conceptual boundaries of agent harnesses, thereby establishing a shared terminological foundation essential for rigorous engineering practice and meaningful academic comparison in intelligent agent systems.
This work addresses the critical yet often overlooked dependence of agent performance on harness engineering, where control logic is typically entangled within code, hindering transferability, reuse, and systematic study. To overcome this limitation, we propose— for the first time—externalizing the high-level control logic of harnesses into editable, executable natural language specifications, supported by a unified Intelligent Harness Runtime (IHR) architecture. The IHR introduces explicit contracts, persistent artifacts, and lightweight adapters to enable modular harness design and cross-task transfer. We validate our approach on programming and computer-use benchmarks, demonstrating its effectiveness through comprehensive experiments. Ablation studies and successful transfers from code-based to text-based harnesses further illustrate the framework’s flexibility and feasibility.
Existing runtime harnesses for programming agents suffer from either oversimplification or excessive complexity, lacking a clear and concise architectural paradigm. This work proposes a harness design centered on the request lifecycle, explicitly delineating three core boundaries: model, execution, and state. By orchestrating a structured sequence—comprising context construction, model decision-making, environmental action, observation feedback, and state continuation—the design enables cross-request state persistence and continual self-improvement through bootstrapping. We implement this paradigm in Coderlet, an open-source prototype system, demonstrating its efficacy in coordinating code generation, environment interaction, and state management. The resulting framework provides a scalable foundation for building high-performance programming agents.
This study systematically investigates how tool architecture influences the behavior and performance of coding agents, holding underlying capabilities constant. Through controlled experiments on repository-scale program repair tasks, six distinct tool interfaces—ranging from bash and structured low-level APIs to natural language search, Python CodeAct, and cognitive scaffolding—are evaluated. Analysis of 11,700 agent trajectories reveals, for the first time, that the architectural design of tools—not merely their functional capacity—plays a critical role: structured low-level interfaces improve consistency across repeated attempts by 4.7×, natural language search increases access to relevant files by over 11%, and CodeAct substantially reduces both action steps (by 41.6%) and token consumption (by 56.3%), whereas cognitive scaffolding yields limited benefits.