Recursive Self-Improvement of Visuomotor Policies through Local Recovery Supervision

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of self-correcting capabilities in existing visuomotor policies, which struggle to autonomously recover from execution failures. To overcome this limitation, we propose a recursive self-improvement framework that achieves closed-loop optimization by offline auditing failure states, leveraging multimodal teacher collaboration to generate local recovery demonstrations, and defining continuous training windows through action-level quality assessment for policy updates. Built upon a BC-RNN architecture, the proposed method significantly enhances policy robustness and generalization. Empirical evaluations demonstrate that it increases the success rate to 88% in the LIBERO-Goal simulation environment and improves performance on the robomimic Can task from 102/130 to 112/130, confirming its effectiveness in enabling autonomous error recovery.
📝 Abstract
Visuomotor policies can execute familiar tasks yet lack the corrective behavior needed after their own mistakes. We present a framework for recursive self-improvement through local recovery supervision. Each round audits the current policy, generates corrective demonstrations at supported failure states, and uses them to update the policy that drives the next round of collection. An offline auditor locates unresolved failures using coarse and dense temporal evidence and specifies observable repair goals. A fixed multimodal agent acts as a tool-using teacher, generating recovery actions through observation, computation, execution, and feedback. The frozen student tests whether each teacher endpoint supports further progress. If continuation fails, the system restores that endpoint and extends the demonstration. Action-level quality assessment then defines continuous training windows with aligned observations, quality weights, and validity masks. Only the student is deployed. In a preliminary LIBERO-Goal study, recovery-augmented post-training achieves 88 successful episodes out of 100 validation scenes, compared with 78 for original-data continuation from the same $\pi_0$ checkpoint. An earlier BC-RNN study on robomimic Can improves success from 102/130 to 112/130 using 26 local recovery segments. Both comparisons match 2,000 additional optimization steps.
Problem

Research questions and friction points this paper is trying to address.

visuomotor policies
error recovery
self-improvement
failure correction
robot learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Recursive Self-Improvement
Local Recovery Supervision
Visuomotor Policies
Multimodal Teacher
Action-level Quality Assessment
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Y
Yuzhi Zhang
Beihang University
X
Xinyu Liu
Beihang University
Yu Zhang
Yu Zhang
Beihang University
surgical planningimage processingmachine learning