Score
Designs and builds agentic AI systems and architectures for single- and multi-agent operation, composing modular components (e.g., initializers, actors, critics, reflectors), defining agent policies, behavior models, and orchestration logic. Simulates and analyzes inter-agent communication and heterogeneous-agent interactions, evaluates agent behavior, safety, human–agent collaboration and test harnesses, and integrates agents into larger systems.
This paper addresses the conceptual conflation between AI Agents and Agentic AI by proposing the first dual-paradigm comparative framework to rigorously distinguish their fundamental differences in design philosophy, autonomy levels, and capability boundaries. Method: It formally defines Agentic AI as requiring four core features—multi-agent collaboration, dynamic task decomposition, persistent memory, and orchestration-level autonomy—thereby extending beyond conventional modular AI Agents. A unified modeling and evaluation stack is developed, integrating LLMs/LIMs, ReAct, RAG, causal modeling, and agent orchestration layers. Contribution: The work establishes a taxonomy spanning architecture, interaction patterns, and autonomy; maps paradigms to canonical applications (e.g., customer service scheduling suits AI Agents, whereas scientific automation and clinical decision support demand Agentic AI); and releases an extensible, interpretable agent development roadmap.
Current AI agents face significant challenges in unifying cognitive modeling, planning, and interactive behavior, as well as ensuring reliable deployment. This paper proposes a unified agent framework that integrates principles from cognitive science, hierarchical reinforcement learning (HRL), and large language model (LLM)-based reasoning to systematically unify perception, decision-making, and interaction. Methodologically, it employs interdisciplinary collaborative modeling, incorporates explainability mechanisms and formal safety constraints, and synergizes multi-agent coordination with deep reinforcement learning to enhance robustness and adaptivity in dynamic, complex environments. The core contributions are threefold: (1) the first end-to-end theoretical pathway bridging cognitive modeling to trustworthy deployment; (2) identification of key technical breakthrough directions; and (3) an architecture blueprint and practical implementation guidelines for next-generation trustworthy, adaptive intelligent systems—balancing theoretical rigor with engineering feasibility.
This study addresses the dual challenges of advancing autonomy, reasoning, and interactive capabilities—while ensuring trustworthiness—in agentic AI systems. Methodologically, it introduces a novel paradigm integrating data-driven learning with structured cognitive modeling, marking the first systematic incorporation of BDI cognitive architectures, multi-agent communication protocols, mechanism design, and institutional modeling from the AAMAS community, augmented with adaptive learning and collaborative reasoning mechanisms. The core contribution is a principled agent framework that balances flexibility, interpretability, and socio-technical embeddability: formal theoretical foundations ensure transparency and accountability of autonomous behavior, while dynamic collaboration mechanisms enable trustworthy human–agent and multi-agent interaction. This framework establishes a unified theoretical pathway and scalable practical foundation for developing sustainable, comprehensible, and governable next-generation autonomous systems.
This paper addresses the conceptual conflation between autonomous AI agents and collaborative multi-agent systems by proposing the first systematic framework for their differentiation. Methodologically, it introduces a structured four-dimensional taxonomy—encompassing planning, memory, coordination, and decision-making—and integrates architectural analysis, paradigmatic taxonomy, protocol modeling, and cross-layer comparison, unifying generative foundation models, tool use, distributed coordination, and memory-augmented techniques. Key contributions include: (1) a rigorous theoretical delineation of the boundary between monolithic agents and emergent collective intelligence; (2) a scalable evolutionary roadmap for agent paradigms; and (3) an empirically grounded agent selection guideline, widely adopted in both industry and academia, which explicitly maps applicability domains and critical bottlenecks of each paradigm—thereby enabling high-reliability research automation and robust design of complex decision-making systems.
This work addresses the fragmentation in current AI agent research stemming from the absence of a systematic architectural framework and unified evaluation standards. To bridge this gap, the paper proposes a comprehensive taxonomy encompassing components, orchestration, and deployment, offering a structured analysis of single- and multi-agent architectures, coordination mechanisms, and application scenarios. It integrates core modules—including large language models, memory systems, world models, planners, tool routers, and critic components—and synthesizes key techniques such as chain-of-thought reasoning, self-reflection, hierarchical planning, and multimodal perception. Building on this foundation, the study consolidates evaluation methodologies—spanning task suites, human preference alignment, and success rates under constraints—elucidates the sources of evaluation complexity, advocates for reproducible benchmarking practices, and highlights critical open challenges in verification, memory management, interpretability, and robustness.
This work addresses the current lack of open-source infrastructure capable of efficiently training and evaluating large-scale agents on complex tasks such as software engineering and computer operation. To this end, we propose a three-service decoupled architecture tailored for agent-environment interaction workloads, which separates the system into three independent services—model, agent, and environment—enabling fine-grained task scheduling, dynamic resource allocation, and unified interface communication. This design allows each component to scale independently and configure resources flexibly, significantly improving training efficiency and resource utilization. Experimental results demonstrate that the system can stably support tens of thousands of concurrent agent tasks, thereby filling a critical gap in infrastructure for large-scale agent training.
This study addresses the challenge of deploying agentic AI in regulated environments, where existing approaches lack a systematic design framework that jointly accounts for autonomy and agency, often failing to balance compliance, auditability, and error correction. The work introduces the first unified model of these two dimensions, defining a two-dimensional hierarchical design space with five operational levels each. It proposes six architectural strategies—checkpoints, escalation mechanisms, multi-agent delegation, tool provisioning, tool sandboxing, and write staging—to enable flexible system configuration under real-world regulatory constraints. Validated through public-sector case studies, the framework establishes a shared terminology and actionable design guidelines, facilitating interpretable, controllable, and compliant AI deployment amid evolving model capabilities and tool fidelity.
This study addresses the lack of systematic investigation into architectural design decisions for non-large language model components in current AI agent systems. The authors propose a protocol-guided, source code–driven empirical analysis method that enables, for the first time, transparent deconstruction of heterogeneous AI agent systems. Through cross-project qualitative coding and co-occurrence analysis of 70 open-source projects, they identify five core design dimensions—sub-agent architecture, context management, tooling systems, security mechanisms, and orchestration—and uncover their combinatorial patterns. Based on these findings, the study further distills five archetypal architectural patterns: lightweight tool-oriented, CLI framework–based, multi-agent orchestrator, enterprise system, and domain-specific vertical architectures.
This work addresses the lack of reliable theoretical foundations for large language model agents in long-horizon, open-ended tasks, where current engineering practices largely rely on empirical trial and error. It systematically introduces classical cybernetics into agent design for the first time, translating its six core principles into actionable design guidelines and proposing a novel “agent cybernetics” framework centered on reliability, sustained operation, and self-improvement. By integrating architectural analysis, failure mode diagnosis, and cross-domain applications—including code generation, computer operation, and automated scientific research—the study identifies critical failure mechanisms and formulates empirically verifiable engineering improvements. This effort establishes both theoretical grounding and practical pathways toward building trustworthy, scalable foundational agents.
Current research on autonomous agent systems often focuses on isolated aspects, lacking a holistic, full-stack framework to guide their design and integration. This work proposes a comprehensive full-stack design framework that spans from foundational large models to multi-agent collaboration, unifying key components including Transformer architectures, efficient fine-tuning techniques (LoRA/MoE), alignment algorithms (RLHF/DPO/GRPO), retrieval-augmented generation (RAG), diverse memory types, MCP and A2A communication protocols, and agent topology structures. Furthermore, it introduces the first taxonomy of agent design patterns. By offering a theoretically grounded yet practically oriented guide—complete with reproducible code, deployment strategies, and evaluation methodologies—this study significantly enhances the constructability, scalability, and real-world effectiveness of intelligent agent systems.