RoboHarn-Evo: Evolving Hierarchical Physical Knowledge for Self-Improving Robotic Manipulation

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the reasoning failures of vision-language models in long-horizon robotic manipulation caused by the absence of local physical interaction feedback. To overcome this limitation, we propose a dual-loop framework that innovatively introduces Hierarchical Physical Knowledge (HPK), transforming repetitive physical interactions into retrievable and correctable structured knowledge assets to enable self-improvement without updating the base model. By integrating knowledge representation, physical feedback revision, and scene grounding techniques, our approach achieves an average success rate improvement of 24.2% on RMBench. Notably, on the GPT-5.5 hold-out set, the success rate increases substantially from 48.3% to 75.0%, while demonstrating significant zero-shot cross-benchmark transferability.
📝 Abstract
Vision-language models can coordinate long-horizon robot manipulation, yet successful task reasoning still depends on whether local physical interactions produce the intended effects. We study how repeated interaction can improve this capability without updating the base model. We introduce RoboHarn-Evo, a dual-loop harness that evolves Hierarchical Physical Knowledge (HPK) from physical experience. HPK couples two levels of reusable knowledge: Task Knowledge captures which subtask should be executed and when it is complete, while Action Knowledge captures object-relative geometric strategies and their physical effects. During execution, the agent retrieves knowledge at the corresponding decision level and grounds it in the current scene under the task goal. Across episodes, physical feedback is used to revise historical knowledge, update its applicability, and organize reusable entries for subsequent retrieval. Experiments on RMBench show that HPK improves average success by up to 24.2 percentage points across different agent models. With 80 interaction rollouts, held-out success rises from 48.3% to 75.0% for GPT-5.5 and from 70.0% to 88.3% for GPT-6. RoboHarn-Evo also resolves over 83% of historical knowledge errors while retaining 95.8% of valid knowledge, and transfers zero-shot from RMBench to RoboDojo with gains of 35.0 and 25.0 percentage points. These results demonstrate that physical interaction can be accumulated into reusable knowledge for improving subsequent manipulation.
Problem

Research questions and friction points this paper is trying to address.

Robotic Manipulation
Vision-Language Models
Physical Interaction
Knowledge Evolution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hierarchical Physical Knowledge
Self-Improving Robotic Manipulation
Dual-loop Harness
Knowledge Evolution
Zero-shot Transfer
S
Shifeng Bao
School of Information, Renmin University of China; Key Laboratory of Data Engineering and Knowledge Engineering, Beijing, China
Fanding Huang
Fanding Huang
Tsinghua University
Semantic SegmentationTest-time AdaptationLarge Language Models
Yihan Lin
Yihan Lin
Assistant Professor, Xiamen University
Brain inspired VisionDeep learningNeuromorphic engineeringComplex networks
Y
Youhe Feng
School of Information, Renmin University of China; Key Laboratory of Data Engineering and Knowledge Engineering, Beijing, China
G
Guanlin Li
School of Information, Renmin University of China; Key Laboratory of Data Engineering and Knowledge Engineering, Beijing, China
C
Chen Zhao
School of Information, Renmin University of China; Key Laboratory of Data Engineering and Knowledge Engineering, Beijing, China
Y
Yang Li
School of Information, Renmin University of China; Key Laboratory of Data Engineering and Knowledge Engineering, Beijing, China
Jiawei He
Jiawei He
XYZ Embodied AI
Computer VisionEmbodied AI
Cheng Chi
Cheng Chi
Columbia University, Stanford University
robotics
Jing Zhang
Jing Zhang
Renmin University of China
large model alignmentmodel compression & inference optimizationdata intelligence