🤖 AI Summary
This study addresses the phenomenon of "incomplete imagination" in world action models under short-horizon control, wherein predictions remain locally plausible yet fail to achieve task objectives. We reveal that this failure stems from short-horizon adaptation rather than deficiencies in the backbone network, and propose Completion-Aware Guidance (CAG), a training-free algorithm. By jointly leveraging visual future prediction and action prediction mechanisms, CAG optimizes the sampling process to bias generated sequences toward task completion. Experimental results demonstrate that CAG reduces the incomplete imagination rate from 79% to 40%, improves the success rate on RoboTwin 2.0 from 64% to 70%, and increases zero-shot simulation success from 69% to 75%.
📝 Abstract
World Action Models (WAMs) predict visual futures and robot actions, yet they remain susceptible to task-incomplete imagination, where plausible, action-consistent predictions omit the transition needed for task completion. In this paper, we show that this failure is not inherent to the world model backbone, but emerges when adapted for short-chunk control, which can repeatedly favor plausible local continuations over task-completing transitions. To address this, we introduce Completion Aware Guidance (CAG), a training-free sampling method that guides generation toward task completion. Across representative WAMs, CAG improves success from 64% to 70% on a RoboTwin 2.0 subset and from 69% to 75% in zero-shot simulation, while reducing task-incomplete imagination from 79% to 40%.