Score
Designs and documents reusable solutions for how agents (people, systems, or devices) exchange actions and feedback, producing interaction models, flows, pattern libraries, frameworks, architectures, and formal interaction specifications. Builds and refines guidelines, detailed interaction behavior, and analytics to map, evaluate, and iterate interaction paradigms and translate high‑level models into implementable interaction specs.
Current AI agents struggle to effectively maintain and leverage user interaction states across multiple devices and over time, leading to insufficient decision coherence. This work proposes a stateful agent architecture that unifies interaction evidence, user-asserted facts, and ongoing requests into a compact, actionable state representation, which is jointly reasoned over with current observations. Built upon multimodal large language models (MLLMs), the architecture implements an end-to-end framework for state management and reasoning and introduces the first cross-device interaction evaluation benchmark. Experiments demonstrate that the proposed approach significantly outperforms four existing agent designs under default settings and across various MLLM variants, confirming the effectiveness and robustness of the introduced state mechanism.
This work proposes a new paradigm for software design tailored to AI agents as primary users, addressing the limitations of traditional human-centric approaches. It formally defines the concept of an “agent interface” for the first time, centering on callable capabilities and emphasizing machine interpretability, composability, and invocation reliability. Guided by the interaction requirements of large language model–based agents, the study introduces a capability-oriented software architecture and corresponding interface specifications. By establishing a conceptual foundation and design framework for AI-native systems, this research advances software engineering beyond monolithic applications toward dynamic, composable ecosystems of interoperable capabilities.
This work addresses the lack of systematic architectural approaches for enterprise-scale multi-agent collaborative systems, particularly in complex scenarios integrating human and AI agents. The authors propose a three-layer design pattern—comprising LLM agents, autonomous agents, and agent communities—that integrates principles from distributed coordination, formal modeling, and a role-protocol-governance structure. For the first time, this framework introduces formal collaboration protocols and role-based governance mechanisms into agent communities, enabling executable specification and verification of organizational, legal, and ethical rules. Validated through a clinical trial matching case study, the architecture demonstrates governable and verifiable human-AI collaboration, offering both formal verification capabilities and actionable design guidance for enterprise deployment.
Current agent system designs often lack grounding in systems theory, resulting in ad hoc architectures prone to hallucination and reasoning flaws that undermine reliability. This work addresses this gap by introducing systems theory into agent architecture design for the first time, proposing a structured framework composed of five core functional subsystems. Building on this foundation, the authors abstract twelve reusable and clearly categorized agent design patterns. Through the reconstruction and validation of representative frameworks such as ReAct, the proposed approach effectively rectifies inherent architectural deficiencies, significantly enhancing modularity, interpretability, and reliability. This contribution establishes a standardized language and a structured development paradigm for agent engineering, offering a principled foundation for future research and practice.
This work addresses the lack of standardized architectural documentation methods tailored to the collaborative and interactive nature of industrial-scale agent-based AI systems, a gap that hinders their maintainability and long-term evolution. To bridge this gap, the paper proposes a specialized modeling approach grounded in agent-oriented architectural styles, extending the C4 layered model with domain-specific modeling vocabularies and viewsets that explicitly capture agents, artifacts, tools, and their coordination patterns. The method further integrates quality-gating mechanisms to ensure documentation consistency. Empirical validation across multiple industrial case studies demonstrates that the proposed approach significantly enhances the clarity and maintainability of architectural documentation while effectively supporting continuous system evolution.
This study systematically investigates how tool architecture influences the behavior and performance of coding agents, holding underlying capabilities constant. Through controlled experiments on repository-scale program repair tasks, six distinct tool interfaces—ranging from bash and structured low-level APIs to natural language search, Python CodeAct, and cognitive scaffolding—are evaluated. Analysis of 11,700 agent trajectories reveals, for the first time, that the architectural design of tools—not merely their functional capacity—plays a critical role: structured low-level interfaces improve consistency across repeated attempts by 4.7×, natural language search increases access to relevant files by over 11%, and CodeAct substantially reduces both action steps (by 41.6%) and token consumption (by 56.3%), whereas cognitive scaffolding yields limited benefits.
This study addresses the current lack of systematic understanding of reusable agent skills in software engineering, particularly regarding their coverage across the software lifecycle. For the first time, it adopts an activity-oriented perspective and conducts a large-scale empirical investigation to collect, categorize, and model software engineering skills from public skill repositories. The work systematically characterizes the types of encapsulated activities, their evolutionary patterns, and evaluation mechanisms. Findings reveal that engineering activities with high contextual dependency are progressively being transformed into reusable skills. Building on these insights, the paper outlines promising future directions, including skill recommendation, structured organization of skill repositories, and enhanced encapsulation strategies for high-context skills, thereby providing both theoretical foundations and practical guidance for agent-driven software engineering.