🤖 AI Summary
This study addresses the challenge in multi-view world action models where implicit geometric relationships hinder the association between scene context and local interactions. By formulating synchronous observations as multi-view projections of the physical world, this work proposes a camera-aware routing mechanism that fuses global states with view-indexed geometry to achieve explicit structured representations. Furthermore, it innovatively introduces epipolar constraints on global states and multi-horizon future depth supervision, which, combined with pretrained video priors, anchor metric scale without requiring depth decoding at inference. The proposed method achieves success rates of 99.1% and 92.07% on the LIBERO and RoboTwin benchmarks, respectively, alongside a 91.3% success rate in real-world robot experiments.
📝 Abstract
World-Action Models (WAMs) couple visual dynamics with action prediction, bringing the rich priors of pretrained video models to robotic manipulation. However, their multi-view interfaces typically tile images or concatenate tokens, leaving the geometric relationships among synchronized cameras implicit. This makes it harder to connect global scene context with the local geometry required for interaction. We introduce the Multi-View Geometry-Aware World-Action Model (MVG-WAM), which organizes these observations as related projections of one physical world rather than separate images on a canvas. Our model combines an epipolar-constrained global state with view-indexed geometric states jointly inferred from synchronized observations. Camera-aware routing supplies each video region with its corresponding geometric context and the shared global state, explicitly structuring the representation used for action prediction. We further ground the geometry-aware representation in metric scale through multi-horizon future-depth supervision, without requiring depth decoding during action rollout. MVG-WAM achieves average success rates of 99.1% on LIBERO and 92.07% on RoboTwin 2.0, demonstrating competitive performance across both benchmarks. Real-world experiments on Cobot Magic further demonstrate a 91.3% success rate across 150 trials spanning three manipulation tasks.