🤖 AI Summary
This work addresses the challenge that existing self-hosted large language model agents (SHCUAs) struggle to meet the stringent requirements of real-time responsiveness, safety, and state continuity when controlling unmanned aerial vehicles via natural language. To overcome this, the paper proposes RT-SHCUA, an architecture that leverages skill contract abstraction to translate semantic outputs into executable skills governed by temporal, stateful, permission-based, fallback, and evidential constraints, thereby decoupling high-level reasoning from onboard control. By integrating edge–onboard collaborative inference, TEE/microcontroller isolation, state consistency verification, and a trusted admission mechanism, RT-SHCUA achieves, for the first time without relocating the large model into a trusted execution environment, auditable, degradable, and state-consistent real-time autonomous control. Prototype evaluation demonstrates that the system meets task response deadlines while enabling trusted access and comprehensive evidence retention of operational behavior.
📝 Abstract
Natural-language control offers a promising interface for unmanned aerial vehicles (UAVs), but directly applying self-hosted computer-use agents (SHCUAs) to UAV control introduces a structural mismatch. SHCUAs are designed for interactive host-side tool use, where delayed agent iterations are often acceptable. UAV control, however, is coupled with continuously changing physical states, strict timing constraints, safety risks, and security accountability. A stale, unauthorized, or tampered agent decision may therefore lead to unsafe or untraceable vehicle behavior.
This paper proposes a real-time and security-oriented restructuring of SHCUA-based UAV control. Instead of allowing an SHCUA to directly issue flight commands, we transform its outputs into contract-bound UAV skill invocations with explicit timing, state, authority, fallback, and evidence semantics. Based on this abstraction, we design an architecture that separates semantic reasoning from onboard execution and security/safety enforcement. Slow cloud or edge reasoning is used for mission understanding, while onboard components validate and dispatch only timely, authorized, and state-consistent skills. Security-critical enforcement points can be protected by TEE-style or microcontroller isolation mechanisms without moving the full language agent or high-frequency flight-control loop into trusted components. Prototype evaluation shows that RT-SHCUA maintains bounded task-level responsiveness while supporting degraded handling, trusted admission, and auditable evidence preservation for SHCUA-mediated UAV actions.