🤖 AI Summary
This study addresses the limitations of existing world-action models, specifically their prediction of redundant visual frames and high inference latency in bimanual robotic manipulation. To this end, it proposes a streaming incremental world-action model that jointly predicts visual deltas and actions rather than generating full frames, coupled with a streaming incremental memory mechanism to substantially reduce the computational overhead of video experts. Architecturally, the framework employs a synergistic multi-architecture design integrating dense anchors, sparse deltas, and action streams. Experimental results demonstrate that the proposed method improves task success rates to 85.4% while reducing training FLOPs by approximately 23% and inference latency by 36%, thereby achieving simultaneous optimization of both manipulation efficiency and computational performance.
📝 Abstract
World-action models (WAMs) transfer visual and motion priors from pretrained video generators to robot control by jointly modeling visual dynamics and actions. Existing WAMs, however, predict dense future frames during training, repeatedly modeling largely unchanged content and coupling action-conditioned dynamics to nuisance appearance variations. At inference, processing each complete observation with the heavy video expert bottlenecks few-step action generation. Accordingly, we propose DeltaWAM, which jointly predicts visual deltas and actions using dense-anchor, sparse-delta, and action streams, with three architectures that differ in representation and computation sharing. We further develop Streaming Delta Memory (SDM), which updates cached anchor context with compact observed deltas, reducing heavy video-expert processing. On RoboTwin, DeltaWAM with SDM improves average success over Fast-WAM from 81.3% to 85.4% in the clean setting and from 75.8% to 83.9% under visual randomization. The three architectures reduce training FLOPs by 17.78-23.77%, while SDM reduces one-step inference latency and FLOPs by 36.57% and 31.55%, respectively; real-world evaluations further show the highest overall success rate and normalized progress among the evaluated policies. Code: https://github.com/AIGeeksGroup/DeltaWAM. Website: https://aigeeksgroup.github.io/DeltaWAM.