iAm.md: Robot Skill Self-Assessment through Agentic Introspection for Unknown Open-Vocabulary Domains

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the failure of behavioral grounding in embodied agents caused by hallucinations or context limitations. To this end, it proposes the iAm.md standard, which anchors deployment evidence using Markdown and introduces open-vocabulary semantic mapping into intermediate representations for the first time. By integrating large language models with visual perception techniques, the method enables code generation-based task execution, introspection, and skill self-assessment. Experiments conducted on TIAGo simulated navigation and manipulation tasks demonstrate that the proposed approach effectively supports agent self-verification and significantly enhances task generalization capabilities in previously unseen scenarios.
📝 Abstract
Agentic AI based on Large Language Model generalization capabilities offers a wide range of potential applications, including planning for embodied tasks. For example, embodied agents based on Foundation models can generate plausible plans in autonomous robotics scenarios. Due to limited context windows or hallucinatory phenomena in the next-token prediction formulation, behaviors may be generated without establishing whether the deployed robot and the observed environment actually support the requested operation, in what we call a "grounding failure". Thanks to the recent improvements in reasoning capabilities of foundation models, autonomous robot behavior generation problem can be formulated as a code generation problem. We present iAm.md, a Markdown standard and generation framework, that allows anchoring this process in complementary forms of deployment evidence. Through open-vocabulary semantic mapping, we combine local vision-language detections and object segmentation and refer them to persistent object records in this intermediate standardized representation, allowing agentic introspection. We then study this new technique on a simulated TIAGo, on navigation-and-manipulation tasks, showing how this standardized representation jointly supports skill self-assessment and executable task generalization.
Problem

Research questions and friction points this paper is trying to address.

grounding failure
embodied agents
hallucination
skill self-assessment
open-vocabulary domains
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic Introspection
Open-Vocabulary Semantic Mapping
Skill Self-Assessment
Grounding Failure
Code Generation
🔎 Similar Papers
No similar papers found.
V
Vincenzo Guarino
Department of Computer, Control and Management Engineering “Antonio Ruberti”, Sapienza University of Rome, Via Ariosto 25, 00185 Rome, Italy
E
Emanuele Musumeci
Department of Computer, Control and Management Engineering “Antonio Ruberti”, Sapienza University of Rome, Via Ariosto 25, 00185 Rome, Italy
Vincenzo Suriani
Vincenzo Suriani
Sapienza University of Rome
Daniele Nardi
Daniele Nardi
Sapienza Univ. Roma, Dept. Computer, Control and Management Engineering
Artificial IntelligenceRoboticsMulti Agent Systems