🤖 AI Summary
This study addresses the failure of behavioral grounding in embodied agents caused by hallucinations or context limitations. To this end, it proposes the iAm.md standard, which anchors deployment evidence using Markdown and introduces open-vocabulary semantic mapping into intermediate representations for the first time. By integrating large language models with visual perception techniques, the method enables code generation-based task execution, introspection, and skill self-assessment. Experiments conducted on TIAGo simulated navigation and manipulation tasks demonstrate that the proposed approach effectively supports agent self-verification and significantly enhances task generalization capabilities in previously unseen scenarios.
📝 Abstract
Agentic AI based on Large Language Model generalization capabilities offers a wide range of potential applications, including planning for embodied tasks. For example, embodied agents based on Foundation models can generate plausible plans in autonomous robotics scenarios. Due to limited context windows or hallucinatory phenomena in the next-token prediction formulation, behaviors may be generated without establishing whether the deployed robot and the observed environment actually support the requested operation, in what we call a "grounding failure". Thanks to the recent improvements in reasoning capabilities of foundation models, autonomous robot behavior generation problem can be formulated as a code generation problem. We present iAm.md, a Markdown standard and generation framework, that allows anchoring this process in complementary forms of deployment evidence. Through open-vocabulary semantic mapping, we combine local vision-language detections and object segmentation and refer them to persistent object records in this intermediate standardized representation, allowing agentic introspection. We then study this new technique on a simulated TIAGo, on navigation-and-manipulation tasks, showing how this standardized representation jointly supports skill self-assessment and executable task generalization.