Score
Designs, builds, and analyzes software systems composed of interacting autonomous agents by specifying agent architectures, behavior models, perception and action interfaces, and environment simulations. Implements and evaluates communication and coordination protocols, decision-making, planning and learning strategies, and mechanisms for negotiation, task allocation, scalability, robustness, and verification to achieve and analyze system-level goals and emergent behaviors.
Current large language model (LLM) agents face challenges in real-world deployment, including inefficiency, error-proneness, and poor maintainability, largely due to their reliance on on-the-fly reasoning and low-level tool invocation. This work introduces, for the first time, a skill-centric agent architecture that formalizes a comprehensive skill lifecycle framework encompassing representation, acquisition, retrieval, and evolution. It positions skills as a complementary mechanism bridging high-level reasoning and operational execution. By integrating key techniques—such as skill representation learning, automated acquisition, semantic retrieval, and continual evolution—and synergizing them with tool use, memory mechanisms, and contextual constraints, the proposed framework establishes a reusable and composable skill system. The paper further surveys representative approaches, open-source resources, and application scenarios, offering both theoretical foundations and practical guidance to enhance the scalability, robustness, and maintainability of intelligent agent systems.
This study addresses the dual challenges of advancing autonomy, reasoning, and interactive capabilities—while ensuring trustworthiness—in agentic AI systems. Methodologically, it introduces a novel paradigm integrating data-driven learning with structured cognitive modeling, marking the first systematic incorporation of BDI cognitive architectures, multi-agent communication protocols, mechanism design, and institutional modeling from the AAMAS community, augmented with adaptive learning and collaborative reasoning mechanisms. The core contribution is a principled agent framework that balances flexibility, interpretability, and socio-technical embeddability: formal theoretical foundations ensure transparency and accountability of autonomous behavior, while dynamic collaboration mechanisms enable trustworthy human–agent and multi-agent interaction. This framework establishes a unified theoretical pathway and scalable practical foundation for developing sustainable, comprehensible, and governable next-generation autonomous systems.
This paper addresses automatic task planning for heterogeneous multi-agent systems operating under dynamic fault conditions. Method: We propose a fault-aware two-layer task synthesis framework. Agent capabilities, operational constraints, and fault models are uniformly encoded using ε-transition-enabled nondeterministic finite automata (NFAs). A decoupled design separates the system model, initial state, and task specification—enabling online injection of fault modes. Formal verification is integrated with heuristic search to ensure solution completeness and optimality while substantially reducing computational overhead. Results: Experiments demonstrate that the approach maintains effectiveness and robustness across diverse fault scenarios. It provides a scalable, formal planning paradigm for high-reliability autonomous decision-making in multi-agent systems.
This work addresses the fragmentation in current AI agent research stemming from the absence of a systematic architectural framework and unified evaluation standards. To bridge this gap, the paper proposes a comprehensive taxonomy encompassing components, orchestration, and deployment, offering a structured analysis of single- and multi-agent architectures, coordination mechanisms, and application scenarios. It integrates core modules—including large language models, memory systems, world models, planners, tool routers, and critic components—and synthesizes key techniques such as chain-of-thought reasoning, self-reflection, hierarchical planning, and multimodal perception. Building on this foundation, the study consolidates evaluation methodologies—spanning task suites, human preference alignment, and success rates under constraints—elucidates the sources of evaluation complexity, advocates for reproducible benchmarking practices, and highlights critical open challenges in verification, memory management, interpretability, and robustness.
This study addresses a critical gap in understanding how task-oriented Agent Plan artifacts in open-source software guide AI-powered coding tools. For the first time, it systematically identifies and analyzes real-world Agent Plan files from open-source projects by screening 36,710 GitHub repositories and conducting qualitative content analysis focused on Markdown-formatted planning documents. The investigation yields 85 valid Agent Plan files that span key engineering activities—including maintenance, design, and implementation—and explicitly articulate task intent while providing concrete execution steps and validation criteria. These findings reveal the instrumental role such plans play in facilitating human-AI collaborative development and underscore their practical value in structuring and communicating software engineering tasks.
This study systematically investigates how tool architecture influences the behavior and performance of coding agents, holding underlying capabilities constant. Through controlled experiments on repository-scale program repair tasks, six distinct tool interfaces—ranging from bash and structured low-level APIs to natural language search, Python CodeAct, and cognitive scaffolding—are evaluated. Analysis of 11,700 agent trajectories reveals, for the first time, that the architectural design of tools—not merely their functional capacity—plays a critical role: structured low-level interfaces improve consistency across repeated attempts by 4.7×, natural language search increases access to relevant files by over 11%, and CodeAct substantially reduces both action steps (by 41.6%) and token consumption (by 56.3%), whereas cognitive scaffolding yields limited benefits.
This work addresses critical challenges in the real-world deployment of large language model (LLM)-driven agent systems, particularly concerning robustness, safety, and reliability. Bridging academic advances with industrial practice, the study presents case studies from software engineering, scientific discovery, and finance to distill reusable design patterns and an evaluation checklist. It integrates key techniques including LLM-based reasoning and planning, multi-agent coordination, validation pipelines, fallback mechanisms, and human-in-the-loop oversight. The proposed cross-domain deployment framework has been validated in pharmaceutical discovery and financial systems, demonstrating significant improvements in stability and trustworthiness of agent systems in real-world settings, thereby narrowing the gap between research innovation and practical implementation.