Direct Experience World-Model Optimization: Learning the World Beyond Action Imitation

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of distribution shift in world models and constrained policy optimization caused by accumulated execution errors during dexterous manipulation. To this end, we propose DEWO, a paradigm that optimizes world representations via visual experience to facilitate action generation. It leverages both successful and failed trajectories for classifier-free guidance and introduces a value head to dynamically activate guidance signals during inference. Notably, this work pioneers post-deployment predictive learning as an active adaptation mechanism that transcends mere action imitation, thereby enhancing the world model's capacity to support control. Experimental results demonstrate that DEWO significantly improves the success rates of three world-action models on DexJoCo tasks. Furthermore, following two rounds of deployment learning on a real robot, the task success rate in specific scenarios increases from 51.0% to 71.7%.
📝 Abstract
World-Action Models (WAMs) couple action generation with predictions of how physical interactions unfold. However, current post-deployment learning paradigms typically improve behavior without requiring better world predictions. Especially in dexterous manipulation, small execution errors can compound in high-dimensional action spaces, hindering policy improvement and pushing interactions beyond the world model's training distribution. Motivated by this, we propose Direct Experience World-Model Optimization (DEWO), a post-deployment learning paradigm for WAMs that, alongside action imitation, refines world representations through visual experience to better condition action generation. Specifically, it identifies interaction turning points and learns from successful and failed futures to support classifier-free guidance. An additional value head estimates task progress from video representations and activates guidance when progress stalls during inference. Across five DexJoCo tasks, DEWO improves average success across all three WAM formulations. Ablations show that visual supervision from successful and failed continuations improves both prediction and control beyond action supervision alone. On four real-world tasks across Wuji and Sharpa, 3 x 3 grid evaluations show that two rounds of deployment learning increase success from 51.0% to 71.7% in cells with at least one initial success, a gain of 20.7 percentage points. These findings support continued predictive learning for improving control through deployment experience, making world modeling an active part of WAM adaptation.
Problem

Research questions and friction points this paper is trying to address.

World-Action Models
Dexterous Manipulation
Post-deployment Learning
Error Compounding
Distribution Shift
Innovation

Methods, ideas, or system contributions that make the work stand out.

World-Action Models
Direct Experience World-Model Optimization
Classifier-Free Guidance
Dexterous Manipulation
Post-Deployment Learning