🤖 AI Summary
This work addresses the lack of natural, context-aware interaction interfaces in existing unmanned aerial systems by introducing OmniAI, an embodied aerial agent that establishes a mobile spatial augmented reality paradigm through adaptive projection switching between an onboard display and environmental surfaces. The system integrates voice commands, mid-air gestures, flight control, and adaptive projection to enable, for the first time, online detection of environmental surfaces without prior mapping—achieved via RGB-D sensing and RANSAC-based plane fitting—and content-adaptive rendering driven by a servo-controlled MEMS laser projector. Context-aware content is further generated through a web-augmented large language model. Experimental results demonstrate functional equivalence between voice and gesture modalities and validate robust, context-sensitive human-drone interaction in dynamic environments.
📝 Abstract
Drones in human environments often lack spatially grounded in- terfaces for situated communication. We present OmniAI, an em- bodied aerial agent that supports surface-adaptive interaction by switching projection between an onboard screen and nearby en- vironmental surfaces. A servo-actuated MEMS laser projector renders text-and-image responses from a web-augmented LLM pipeline. Projection surfaces are detected online using RGB-D sensing and RANSAC plane fitting, without pre-mapped geometry. OmniAI provides functionally equivalent voice and gesture con- trol for both drone motion and projected content. By combining speech, mid-air gestures, adaptive projection, and aerial mobility, OmniAI demonstrates a mobile spatial AR interface for context- aware human-drone interaction.