🤖 AI Summary
This study addresses the challenge of translating natural language intents into robot-executable actions within dynamic, unknown environments by proposing an intent-driven dual-AI collaborative framework. The framework leverages large language models to generate constrained executable code and integrates vision-language models for semantic grounding. Its core innovation lies in an adaptive replanning mechanism triggered by geometric and semantic thresholds, which achieves closed-loop control through runtime monitoring. Experimental results demonstrate that the proposed approach robustly executes complex instructions under bounded reaction cycles, enabling effective robot control in dynamic settings. These findings validate the reliability and generalization capability of generative AI for real-time embodied intelligence tasks.
📝 Abstract
Deploying robots as Complex Adaptive Systems (CAS) in unknown and dynamic environments necessitates a transition from rigid command libraries toward intention-based autonomy, as natural language represents the only medium capable of articulating complex goals beyond the capacity of finite instruction sets. While Large Language Models (LLMs) offer a path toward natural language goal description, their integration introduces significant challenges: the formalization gap between imprecise intentions and executable actions, the taxonomy gap induced by unpredictable environments, and the challenge of maintaining temporal state and progress awareness. This work introduces an architectural framework that enables robotic control by leveraging generative AI. The system follows a dual-AI design: an LLM translates high-level intentions into executable program code restricted to a formal robotic library and constrained by verifiable syntax, while a Vision-Language Model (VLM) provides semantic grounding via a distillation process. To ensure robustness, the framework incorporates environment-driven replanning triggers based on geometric and semantic thresholds, complemented by continuous runtime monitoring and an adaptive planning loop. Benchmarked across frontier models, our framework architecture demonstrates that grounding generative AI in a reactive, constrained loop enables robust fulfillment of complex intentions in dynamic and unknown environments.