🤖 AI Summary
This study addresses the reliability challenges faced by LLM agents executing tasks across heterogeneous devices, where failures arise from unachieved effects, uncertain outcomes, and shifting preconditions. To mitigate these issues, this work proposes the Device Capability Contract (DCC) architecture. Using DCCs as a semantic foundation, the approach unifies invocation conditions and recovery rules for planning and execution. By integrating persistent state management, observation-based recovery mechanisms, and cross-framework runtime verification, it formalizes the execution lifecycle and establishes soundness properties governing completion and recovery authorization. Experimental results demonstrate that the proposed architecture significantly reduces spurious completions and redundant operations, facilitates necessary state repairs, and prevents invalid invocations, thereby effectively enhancing the reliability of task execution across multiple domains.
📝 Abstract
Agents based on large language models (LLMs) can access heterogeneous devices through tools and APIs, but reliable execution must account for unmet effects, uncertain outcomes, and changing prerequisites. A command may be acknowledged without producing its intended effect, while missing feedback may obscure an action that has already succeeded. We present Agent Device Foundation--Execution Assurance (ADF-EA), an architecture that connects agent planning and device execution through shared capability contracts. Device Capability Contracts (DCCs) unify invocation conditions, intended effects, evidence requirements, and recovery rules across heterogeneous interfaces. Agents use these contracts to plan, while the runtime applies the same semantics to authorize actions, verify effects, and govern continuation and completion. Persistent execution state retains verified progress, unresolved outcomes, and remaining budgets across plan revisions, enabling observation-based recovery, authorized retries, and necessary state repair. We formalize the execution lifecycle and establish conditional soundness properties for completion and recovery authorization. Evaluations span multiple LLMs, five agent frameworks, and simulated process-control, household, and robotic manipulation domains. Compared with direct invocation and existing execution-checking approaches, ADF-EA reduces false completion and unnecessary repetition, supports necessary state repair, prevents calls to unavailable capabilities, and preserves permitted task completion and recovery. These results demonstrate DCCs as a reusable semantic foundation for agent autonomy across heterogeneous devices, unifying capability-based planning, evidence-grounded execution, and authorized recovery within one architecture.