🤖 AI Summary
This study addresses the reasoning failures of vision-language models in long-horizon robotic manipulation caused by the absence of local physical interaction feedback. To overcome this limitation, we propose a dual-loop framework that innovatively introduces Hierarchical Physical Knowledge (HPK), transforming repetitive physical interactions into retrievable and correctable structured knowledge assets to enable self-improvement without updating the base model. By integrating knowledge representation, physical feedback revision, and scene grounding techniques, our approach achieves an average success rate improvement of 24.2% on RMBench. Notably, on the GPT-5.5 hold-out set, the success rate increases substantially from 48.3% to 75.0%, while demonstrating significant zero-shot cross-benchmark transferability.
📝 Abstract
Vision-language models can coordinate long-horizon robot manipulation, yet successful task reasoning still depends on whether local physical interactions produce the intended effects. We study how repeated interaction can improve this capability without updating the base model. We introduce RoboHarn-Evo, a dual-loop harness that evolves Hierarchical Physical Knowledge (HPK) from physical experience. HPK couples two levels of reusable knowledge: Task Knowledge captures which subtask should be executed and when it is complete, while Action Knowledge captures object-relative geometric strategies and their physical effects. During execution, the agent retrieves knowledge at the corresponding decision level and grounds it in the current scene under the task goal. Across episodes, physical feedback is used to revise historical knowledge, update its applicability, and organize reusable entries for subsequent retrieval. Experiments on RMBench show that HPK improves average success by up to 24.2 percentage points across different agent models. With 80 interaction rollouts, held-out success rises from 48.3% to 75.0% for GPT-5.5 and from 70.0% to 88.3% for GPT-6. RoboHarn-Evo also resolves over 83% of historical knowledge errors while retaining 95.8% of valid knowledge, and transfers zero-shot from RMBench to RoboDojo with gains of 35.0 and 25.0 percentage points. These results demonstrate that physical interaction can be accumulated into reusable knowledge for improving subsequent manipulation.