DeltaWAM: Change-Centric Visual Foresight via Delta Tokens for an Efficient World-Action Model

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the computational redundancy and high costs associated with full reconstruction in pixel-level world models by proposing a change-centric visual prediction framework. The core innovation lies in introducing the Delta Token as the fundamental prediction unit, which compactly encodes DINO feature differences between consecutive frames to replace full-scene reconstruction. By integrating a latent-space autoregressive mechanism with a flow-matching action expert, the framework generates robot manipulation sequences. Evaluated on the LIBERO benchmark, the proposed method achieves an average success rate of 92.8% while demonstrating strong generalization capabilities. Furthermore, it attains remarkable spatiotemporal efficiency, requiring only 142.1 milliseconds per inference step with a peak memory consumption of 3.86 GB, significantly outperforming conventional full-reconstruction approaches.
📝 Abstract
World-Action Models (WAMs) offer visual foresight for robotic manipulation, but pixel-space models repeatedly reconstruct entire future scenes, incurring high computational cost and spatio-temporal redundancy. In physical manipulation, consecutive frames often share most of their visual context; the changes between them are what an action policy needs to anticipate. We introduce DeltaWAM, a change-centric WAM that makes a compact delta token the unit of future prediction. Each token is a single vector encoding changes between consecutive dense DINO feature maps. DeltaWAM builds on DeltaWorld, a latent world model pretrained on large-scale videos, to autoregressively predict one delta token per future frame. A flow-matching action expert then conditions on the predicted transitions and current DINO features, which serve as spatial anchors, to generate action chunks. Trained for 256 GPU hours on two H100 GPUs, DeltaWAM has 0.725B parameters and achieves 92.8% average success on LIBERO. It also shows robust generalization under procedural perturbations on LIBERO-Pro. Inference takes 142.1 ms per action chunk with 3.86 GB peak memory.
Problem

Research questions and friction points this paper is trying to address.

World-Action Models
visual foresight
computational cost
spatio-temporal redundancy
robotic manipulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Delta Token
World-Action Model
Flow-Matching Action Expert
Visual Foresight
Latent World Model
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.