🤖 AI Summary
This study addresses the inefficiency in world action models caused by supervision signals dominated by appearance redundancy rather than action dynamics. To this end, we propose an action tangent field reconstruction supervision mechanism that leverages residual VAE representations and action-conditioned world model probing to distill local first-order structures into the model, guiding the denoising process to focus on action-coupled dynamic changes. Furthermore, multi-step VideoDiT single-step distillation is incorporated to optimize inference efficiency. Experiments demonstrate that our method significantly enhances robustness against illumination and background perturbations on benchmarks such as LIBERO, outperforming existing strategies without requiring large-scale pretraining. Ultimately, this work facilitates a paradigm shift from appearance-dominated prediction toward action-driven dynamic learning.
📝 Abstract
World Action Models (WAM) improve robot policies by augmenting sparse action supervision with dense future prediction. However, much of the predictable future is dominated by appearance and scene persistence rather than action-dependent dynamics. We observe that several recent WAM designs, including optical flow, motion-centric representations, and latent actions, can be understood from a common perspective in which world supervision becomes more efficient as it contains a higher proportion of action-relevant variation. Based on this insight, we introduce Action Tangent Fields, which reformulate world supervision through a local Taylor expansion of how actions induce changes in future dynamics. We represent future dynamics in Residual-VAE space, where the future latent remains recoverable from the current latent and its residual, and use a strong action-conditioned world model (ACWM) to probe the local correspondence between action variations and residual-world variations. This local first-order structure is distilled into the WAM to guide its denoising supervision toward dynamics that are more tightly coupled to action, rather than merely predictable from appearance. Across LIBERO-Plus, RoboTwin, and RoboTwin2.0-Plus, our method consistently improves robustness to lighting, background, camera, layout, and other environmental perturbations. Despite using no large-scale embodied pretraining, it achieves stronger robustness under several distribution shifts than pretrained policies. We further distill multi-step VideoDiT denoising into a single step for efficient inference. Our results suggest that effective WAM supervision should remain information-rich while concentrating its predictive capacity on the directions along which actions change the future.