MVG-WAM: Multiple View Geometry-Aware World-Action Modeling for Robotic Manipulation

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge in multi-view world action models where implicit geometric relationships hinder the association between scene context and local interactions. By formulating synchronous observations as multi-view projections of the physical world, this work proposes a camera-aware routing mechanism that fuses global states with view-indexed geometry to achieve explicit structured representations. Furthermore, it innovatively introduces epipolar constraints on global states and multi-horizon future depth supervision, which, combined with pretrained video priors, anchor metric scale without requiring depth decoding at inference. The proposed method achieves success rates of 99.1% and 92.07% on the LIBERO and RoboTwin benchmarks, respectively, alongside a 91.3% success rate in real-world robot experiments.
📝 Abstract
World-Action Models (WAMs) couple visual dynamics with action prediction, bringing the rich priors of pretrained video models to robotic manipulation. However, their multi-view interfaces typically tile images or concatenate tokens, leaving the geometric relationships among synchronized cameras implicit. This makes it harder to connect global scene context with the local geometry required for interaction. We introduce the Multi-View Geometry-Aware World-Action Model (MVG-WAM), which organizes these observations as related projections of one physical world rather than separate images on a canvas. Our model combines an epipolar-constrained global state with view-indexed geometric states jointly inferred from synchronized observations. Camera-aware routing supplies each video region with its corresponding geometric context and the shared global state, explicitly structuring the representation used for action prediction. We further ground the geometry-aware representation in metric scale through multi-horizon future-depth supervision, without requiring depth decoding during action rollout. MVG-WAM achieves average success rates of 99.1% on LIBERO and 92.07% on RoboTwin 2.0, demonstrating competitive performance across both benchmarks. Real-world experiments on Cobot Magic further demonstrate a 91.3% success rate across 150 trials spanning three manipulation tasks.
Problem

Research questions and friction points this paper is trying to address.

World-Action Models
Multi-view Geometry
Robotic Manipulation
Visual Dynamics
Innovation

Methods, ideas, or system contributions that make the work stand out.

World-Action Model
Multi-View Geometry
Epipolar Constraint
Camera-Aware Routing
Future-Depth Supervision
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
W
Wenbo Chen
The Hong Kong University of Science and Technology (Guangzhou)
T
Tianfu Li
The Hong Kong University of Science and Technology (Guangzhou)
Haoxuan Xu
Haoxuan Xu
Beihang University
computer vision
Z
Zhihao Cao
ETH Zurich
Z
Zhenghan Chen
Zhejiang University
Z
Zhengming Zhu
EPFL
Z
Zizhou Luo
University of Zurich
G
Guosheng Yang
The Hong Kong University of Science and Technology (Guangzhou)
Y
Yuan Liu
The Hong Kong University of Science and Technology
Lujia Wang
Lujia Wang
Hong Kong University of Science and Technology
Cloud roboticsLifelong federated learningresource/Task allocation for cloud-edge systems and applications for autonomous dri
Wen Chen
Wen Chen
PhD, The Chinese University of Hong Kong
Point Cloud RegistrationSLAMState Estimation
Haoang Li
Haoang Li
Assistant Professor, Hong Kong University of Science and Technology (Guangzhou)
Robotics3D Computer Vision