SEES: A Self-Evolving Embodied System via Failure-Guided VLA Policy Adaptation

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the reliability limitations and dependence on expert demonstrations in Vision-Language-Action (VLA) policies for long-horizon tasks, which arise from data bias. To this end, we propose a failure-guided self-evolving embodied system. Methodologically, long-horizon tasks are decomposed into atomic skills with dynamic routing. By monitoring failure bottlenecks, large language models automatically generate simulation environments and evaluation criteria, while online reinforcement learning updates adapters to achieve policy optimization without expert intervention. Experimental results demonstrate that our approach significantly enhances the long-horizon performance of diverse VLA backbones, exhibiting continuous improvement and strong generalization capabilities on unseen tasks.
📝 Abstract
Recent vision-language-action (VLA) policies demonstrate promising generalization across diverse short-horizon tasks. However, they remain unreliable on long-horizon tasks, partly because the large-scale training data is biased toward single-stage manipulation tasks that are cheaper to demonstrate. A single weak atomic skill can cause failures across multiple multi-stage tasks. To address such failures, existing methods often require experts to identify the bottleneck and provide additional demonstrations, making the improvement costly and potentially impractical after deployment. To this end, we present a Self-Evolving Embodied System (SEES) that learns from failures and improves the VLA policy without additional expert demonstrations. SEES decomposes long-horizon tasks into atomic tasks and routes them to corresponding family policies. Each family consists of related atomic skills that share one VLA adapter. During execution, the system automatically monitors atomic-task outcomes to identify the most frequently failing atomic skills as the current bottlenecks. To overcome these bottlenecks, SEES constructs tailored RL tasks in simulation by restoring previously encountered states and generating task-specific success criteria with an LLM. Online RL updates the shared family adapters to promote positive transfer among related atomic skills and cumulative improvement across evolution rounds. Extensive experiments show that SEES can be integrated with different VLA backbones to progressively improve their long-horizon performance. We also observe continued improvement on unseen tasks, providing evidence of transfer beyond the evolution settings.
Problem

Research questions and friction points this paper is trying to address.

vision-language-action policy
long-horizon tasks
failure-guided adaptation
embodied system
self-evolution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Evolving Embodied System
Vision-Language-Action Policy
Failure-Guided Adaptation
Reinforcement Learning
Long-Horizon Tasks
🔎 Similar Papers
No similar papers found.
Z
Ziwen Li
Mohamed bin Zayed University of Artificial Intelligence
H
Hanlue Zhang
Mohamed bin Zayed University of Artificial Intelligence
Z
Zhenyang Ren
University of Sydney
T
Tianyu Huang
University of Sydney
Runqi Lin
Runqi Lin
PhD student, The University of Sydney
Machine LearningAI SafetyTrustworthy MLAdversarial Robustness
H
Haoyu Wang
Mohamed bin Zayed University of Artificial Intelligence
Z
Zhengqing Gao
Mohamed bin Zayed University of Artificial Intelligence
Y
Yandong Guo
AI2 Robotics
F
Fakhri Karray
University of Waterloo, Mohamed bin Zayed University of Artificial Intelligence
Tongliang Liu
Tongliang Liu
Director, Sydney AI Centre, University of Sydney & Mohamed bin Zayed University of AI
Machine LearningLearning with Noisy LabelsTrustworthy Machine Learning
Chris Russell
Chris Russell
Associate Professor, University of Oxford
Ethical Machine LearningComputer VisionOptimisationEthical AI
Mingming Gong
Mingming Gong
University of Melbourne & Mohamed bin Zayed University of Artificial Intelligence
Causal InferenceMachine LearningComputer Vision