🤖 AI Summary
General-purpose vision-language models (VLMs) lack embodied cognition regarding robotic body structures and action effects. To address this limitation, this work proposes KnowBody, a framework that explicitly models action-body relationships as a queryable and revisable knowledge graph to guide action planning with frozen model weights. This approach enables self-improving control without fine-tuning by continuously updating body knowledge from online interaction feedback and verifying historical dependencies. Experimental results demonstrate that the task success rate increases from 25% to 75%, while the number of planning rounds required for successful trials is significantly reduced. Furthermore, knowledge reuse improves overall planning efficiency by 29% to 53%.
📝 Abstract
A general-purpose vision-language model can understand a task goal without knowing how a particular robot's motion and functional parts produce the intended effect. We introduce KnowBody, a harness that makes these action-relevant body relations explicit, queryable, and revisable while keeping the model weights frozen. Initialized from one off-task trajectory, a partial body model guides action selection and the interpretation of past interactions. New evidence refines the model, and knowledge dependent on revised body estimates is rechecked before reuse. Across 32 fixed-budget trials on four real-robot tasks, initialized KnowBody achieves 75% completion versus 25% for the native harness and requires fewer planner rounds on successful trials in tasks completed by both. With persistent updates enabled, planner rounds decrease by 29-53% from the first to the fifth recorded success.