Score
Designs, builds, or analyzes a modular agent observation interface (perception layer, AOI) that ingests and synchronizes continuous sensory inputs and telemetry, preprocesses and converts them into structured representations usable by downstream decision or control modules. This competence includes implementing buffering and rate-control to decouple observation flow from action execution, gating and noise-filtering, and facilities to persist, export, or query multimodal observations (for example, visual/audio streams or textified transcripts).
Existing agents are constrained by low-frequency screenshot observations, rendering them unable to perceive dynamic information such as video, audio, and transient UI events. This work proposes the Agent-Computer Observation Interface (AOI), which treats the observation interface as an independent design dimension for the first time. By decoupling continuous adaptive perception from discrete actions, AOI establishes a model-agnostic perceptual layer. It employs a three-component gating mechanism—keyframe capture, volume-triggered speech transcription, and visual narrative textualization—to enable efficient perception of dynamic computing environments. Evaluated on DynaCU-Bench, AOI boosts performance across diverse computer-use (CU) models by 17–48 percentage points without any retraining, elevating audio task success rates from near 0% to 100%.
To address the observability gap arising from non-deterministic behaviors of LLM-driven AI agents, this paper proposes the first observability framework integrating process mining, causal discovery, and LLM static analysis. Methodologically: (1) process discovery is performed on execution trace logs to model behavioral variation patterns; (2) causal inference identifies root causes of such variations; and (3) static semantic analysis of prompts and code—conducted via LLMs—distinguishes intentional variations (e.g., strategic adjustments) from unintentional ones (e.g., hallucinations or logical drift). Our key contribution is the first synergistic application of process mining and causal discovery for fine-grained, interpretable attribution of AI agent behavior variations. The framework enables developers to locate ambiguous specifications and diagnose unintended execution branches, thereby significantly improving debugging efficiency and system controllability.
To address the inflexibility in modeling, difficulty in verification, and error-proneness in implementation of interaction protocols in multi-agent systems, this paper proposes a communication-protocol-based, interaction-oriented programming framework. The framework centers on formal protocol models to precisely specify inter-agent interaction behaviors; integrates a protocol verifier and agent middleware to enable formal verification of safety and liveness properties, as well as automated code generation; and establishes a complete toolchain spanning protocol modeling, verification, code generation, and execution. Compared with conventional approaches, our framework significantly reduces protocol design errors—by approximately 62% in empirical evaluation—while improving development efficiency and system maintainability. It provides a verifiable, engineering-ready infrastructure for trustworthy multi-agent collaboration.
This work proposes a “layered attribution” diagnostic framework to disentangle the origins of inscrutable behaviors exhibited by AI agents in complex social systems, which are often conflated between internal representations and external constraints. The framework systematically distinguishes a foundational computational layer—encompassing architecture, memory, and perception—from a behavioral modulation layer comprising identity, goals, social interactions, and institutional constraints, thereby integrating representation learning, multi-agent modeling, and institutional analysis into a unified two-tier diagnostic architecture. It yields three key insights: behavioral substitutability validity hinges on the coupling among model, task, and layer; human–AI behavioral discrepancies can serve as diagnostic signals; and effective governance presupposes precise source attribution. This approach establishes a theoretical foundation for interpreting and governing AI behavior.
Current distributed agent systems lack a unified, implementation-agnostic runtime architecture, making it difficult to effectively govern intent, permissions, uncertainty, behavioral coordination, and traceability. This work proposes an Agent Operating System (AOS), which introduces the first dual-plane reference architecture for distributed agent systems: a control and governance plane responsible for policy enforcement, auditing, and human oversight, and a runtime and coordination plane handling workflow execution, model routing, memory coordination, and scheduling. By standardizing interfaces that decouple governance from runtime responsibilities, AOS enables flexible composition of heterogeneous components and formalizes core concepts, interface objects, and deployment configurations. This architectural foundation supports the development of trustworthy, auditable, scalable, and interoperable agent systems, while also outlining key research challenges in the field.
This work addresses key challenges in complex industrial settings—namely, poor scalability, limited observability, and difficulties in autonomous evolution—faced by multi-agent systems. To overcome these limitations, the authors propose OxyGent, a novel framework that introduces a unified Oxy abstraction to encapsulate agents, tools, large language models, and reasoning pipelines as plug-and-play atomic components. This design enables LEGO-like composition, non-intrusive monitoring, and continuous system evolution. Core innovations include a permission-driven dynamic planning mechanism that replaces rigid workflows, automatic execution graph generation, and the OxyBank platform, which facilitates automatic feedback, annotation, and co-evolution of AI assets. Empirical results demonstrate that OxyGent substantially enhances the robustness, flexibility, and scalability of multi-agent systems.
This work addresses the challenge of effectively supervising complex, long-running AI agents—a task often beyond human capability—by introducing AgentGUI, a locally deployed graphical user interface that enables real-time observation and intervention in multi-agent systems. AgentGUI is the first system to integrate trajectory visualization with both manual and automated control mechanisms, while maintaining compatibility with mainstream open-source and state-of-the-art agent frameworks to support unified human-AI collaborative oversight. Experimental results demonstrate that users leveraging AgentGUI identify critical trajectory elements 38% faster (p = 0.023), and task completion rates for smaller models improve by up to 34 percentage points.
Current AI agents struggle to effectively maintain and leverage user interaction states across multiple devices and over time, leading to insufficient decision coherence. This work proposes a stateful agent architecture that unifies interaction evidence, user-asserted facts, and ongoing requests into a compact, actionable state representation, which is jointly reasoned over with current observations. Built upon multimodal large language models (MLLMs), the architecture implements an end-to-end framework for state management and reasoning and introduces the first cross-device interaction evaluation benchmark. Experiments demonstrate that the proposed approach significantly outperforms four existing agent designs under default settings and across various MLLM variants, confirming the effectiveness and robustness of the introduced state mechanism.