RoboHarn-Evo: Evolving Hierarchical Physical Knowledge for Self-Improving Robotic Manipulation
This study addresses the reasoning failures of vision-language models in long-horizon robotic manipulation caused by the absence of local physical interaction feedback. To overcome this limitation, we propose a dual-loop framework that innovatively introduces Hierarchical Physical Knowledge (HPK), transforming repetitive physical interactions into retrievable and correctable structured knowledge assets to enable self-improvement without updating the base model. By integrating knowledge representation, physical feedback revision, and scene grounding techniques, our approach achieves an average success rate improvement of 24.2% on RMBench. Notably, on the GPT-5.5 hold-out set, the success rate increases substantially from 48.3% to 75.0%, while demonstrating significant zero-shot cross-benchmark transferability.