EVO-WAM: Evolving World Action Models through Video-Action Verification

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the difficulty of adapting robotic policies to novel tasks without expert demonstrations, where inconsistencies between videos generated by world action models and corresponding actions frequently lead to execution failures. To this end, this work proposes a pioneering self-evolutionary framework that operates without external feedback. By generating video-action trajectories autonomously for unsupervised adaptation, the method integrates state prediction with multi-frame context anchoring. Furthermore, it employs vision-language models (VLMs) to evaluate task completion and inverse dynamics models to verify action consistency, iteratively enhancing policy capabilities through training. Empirical evaluations demonstrate that the proposed approach improves success rates by approximately 2.5× across seven unseen tasks on RoboTwin 2.0, while significantly increasing performance on real-world long-horizon composite tasks from 20% to 76.7%.
📝 Abstract
Improving robot policies on new tasks without collecting additional expert demonstrations remains a central challenge in robot learning. World action models (WAMs) use broad video priors to jointly predict future videos and actions, offering a potential source of supervision for adapting to new tasks. However, generated videos may fail to depict task completion, and even visually successful videos may be paired with inconsistent actions that lead to execution failure. We propose EVO-WAM, a framework that adapts WAMs to unseen tasks by learning from their own generated video-action trajectories, without executing candidate actions in an external environment. First, we augment WAM training with state prediction and anchored multi-frame context to enable complete autoregressive rollouts without external execution feedback. Second, we identify reliable training experience by selecting task-completing prefixes with a vision-language model and verifying their video-action consistency with an inverse dynamics model. Third, we iteratively train the WAM on verified prefixes and generate new rollouts with the updated model. On seven unseen RoboTwin 2.0 tasks, EVO-WAM increases average success rates from 26.9% to 68.0% for Cosmos3 and from 28.5% to 46.4% for DreamZero, reaching approximately $2.5\times$ and $1.6\times$ their initial success rates. On three unseen long-horizon composite tasks in the real world, it improves Cosmos3's average success rate from 20.0% to 76.7%, a gain of 56.7 percentage points. Project Page: https://evo-wam.github.io/.
Problem

Research questions and friction points this paper is trying to address.

World Action Models
Robot Learning
Task Adaptation
Video-Action Consistency
Self-supervised Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

World Action Models
Video-Action Verification
Self-Evolution
Inverse Dynamics Model
Robot Learning
🔎 Similar Papers
No similar papers found.