Score
Designs, implements, and evaluates software components, interfaces, and frameworks that enable autonomous agents to discover, invoke, compose, and manage external tools, APIs, and resources; this includes adapters, action managers, orchestration logic, instrumentation, and debugging utilities. It also covers analysis of agent–tool interactions for correctness, reliability, performance, and safety (e.g., sandboxing, access control, and failure handling).
This work addresses the limitations of large language model agents in reliably executing complex tasks due to inherent cognitive burdens that cannot be overcome by parametric capabilities alone. The authors propose a system-level framework centered on externalization, treating memory, skills, and protocols as three interdependent cognitive artifacts coordinated through a unified execution architecture. Grounded in cognitive artifact theory, the framework integrates modular external components with runtime coordination mechanisms, offering a cohesive analytical lens for agent infrastructure. The study demonstrates that performance gains increasingly stem from external cognitive infrastructure rather than model scaling alone, highlighting promising directions such as self-evolving architectures and shared infrastructures, while also identifying associated challenges in evaluation and governance.
This study systematically investigates how tool architecture influences the behavior and performance of coding agents, holding underlying capabilities constant. Through controlled experiments on repository-scale program repair tasks, six distinct tool interfaces—ranging from bash and structured low-level APIs to natural language search, Python CodeAct, and cognitive scaffolding—are evaluated. Analysis of 11,700 agent trajectories reveals, for the first time, that the architectural design of tools—not merely their functional capacity—plays a critical role: structured low-level interfaces improve consistency across repeated attempts by 4.7×, natural language search increases access to relevant files by over 11%, and CodeAct substantially reduces both action steps (by 41.6%) and token consumption (by 56.3%), whereas cognitive scaffolding yields limited benefits.
研究解决了工具调用与操作性之间的差距问题,通过Agent-First Tooling机制和AFT-Bench框架来确保代理在不确定性下安全继续执行。
Current LLM-driven autonomous agents lack scalable, secure, and maintainable architectures for tool orchestration, hindering large-scale deployment. To address this, we propose the novel “Control Plane as a Tool” paradigm, which— for the first time—abstracts tool scheduling, security policies, and extensibility mechanisms into a unified, pluggable tool interface, thereby decoupling control logic from the LLM agent core. Leveraging modular routing protocols and production-grade encapsulation, our approach significantly reduces integration complexity while enabling dynamic tool registration/removal and runtime policy updates. Empirical evaluation across diverse application scenarios demonstrates over 40% improvement in both horizontal scalability efficiency and security robustness. The architecture provides a reusable, evolution-aware foundation for controllable, embodied intelligent agents.
This study addresses the lack of systematic investigation into architectural design decisions for non-large language model components in current AI agent systems. The authors propose a protocol-guided, source code–driven empirical analysis method that enables, for the first time, transparent deconstruction of heterogeneous AI agent systems. Through cross-project qualitative coding and co-occurrence analysis of 70 open-source projects, they identify five core design dimensions—sub-agent architecture, context management, tooling systems, security mechanisms, and orchestration—and uncover their combinatorial patterns. Based on these findings, the study further distills five archetypal architectural patterns: lightweight tool-oriented, CLI framework–based, multi-agent orchestrator, enterprise system, and domain-specific vertical architectures.
AI coding assistants and autonomous agents are becoming integral to software development workflows, reshaping how code is produced, reviewed, and maintained. While recent research has focused mainly on the capabilities and impacts of productivity of these systems, much less attention has been paid to accountability: who is responsible when agents generate, modify, or recommend code? In practice, accountability is defined through the Terms of Service (ToS) and related policy documents that govern the use of AI-powered development tools. In this vision paper, we present a comparative analysis of the Terms of Service for widely used AI coding assistants and agent-enabled development tools. We examine how these documents allocate ownership, responsibility, liability, and disclosure obligations between tool providers and software developers, and we identify common patterns and divergences between providers. Our analysis reveals a consistent tendency to shift responsibility for correctness, safety, and legal compliance onto users, as well as substantial variation in how providers address issues such as indemnification, data reuse, and acceptable use. Based on these findings, we argue that existing policy frameworks are poorly aligned with increasingly agent-mediated and autonomous software development workflows. We outline a research roadmap for accountable agents in software engineering, identifying challenges and opportunities for modeling responsibility, designing governance artifacts, developing tooling that supports accountability, and conducting empirical studies of developers' perceptions and practices.
This study addresses a critical gap in understanding how task-oriented Agent Plan artifacts in open-source software guide AI-powered coding tools. For the first time, it systematically identifies and analyzes real-world Agent Plan files from open-source projects by screening 36,710 GitHub repositories and conducting qualitative content analysis focused on Markdown-formatted planning documents. The investigation yields 85 valid Agent Plan files that span key engineering activities—including maintenance, design, and implementation—and explicitly articulate task intent while providing concrete execution steps and validation criteria. These findings reveal the instrumental role such plans play in facilitating human-AI collaborative development and underscore their practical value in structuring and communicating software engineering tasks.
This study addresses the behavioral divergence of AI agents and the inadequacy of conventional testing in ensuring their reliability by proposing a programming-language perspective that treats agents as programmable artifacts, systematically restructuring their reliability framework. Methodologically, rather than pursuing deterministic execution, this work adopts a structured management paradigm. It employs trajectory- and state-based behavioral specifications to define expected outcomes, integrating pre-deployment static analysis with runtime dynamic monitoring for full-lifecycle governance. Consequently, this approach renders agent behavior both reason-able and controllable while supporting continuous iterative repair from operational failures. Ultimately, this research offers a novel pathway for constructing highly reliable AI agent systems.
研究解决了AI代理与工具间的工作流失败问题,通过提出一种效应历史模型和异常目录,并探讨了现有工具接口的局限性。
论文提出Agent-Integrated Software模式和Intent-Level Interaction Abstraction方法来解决智能代理与现有应用程序集成时的协调问题,确保任务级交互与应用行为的一致性。
This work proposes a software engineering–inspired approach to enhance the controllability and engineering rigor of large language model (LLM) agents by treating agent skills as modular software components. For the first time, core software engineering principles—including single responsibility, separation of interface and implementation, low coupling, and token economy—are systematically applied to guide skill design. The authors establish a behavior-evaluation-driven skill development pipeline integrating UML modeling, phased loading mechanisms, and standardized skill descriptions, while formally defining skill structure and loading models. The study further identifies canonical implementation patterns and anti-patterns, and formulates decision rules for selecting among coordination mechanisms such as memory integration and sub-agent delegation. This framework provides developers with actionable guidelines for building reusable, maintainable skills and offers criteria for evaluating trustworthiness in third-party skills.