ACG-WAM: World-Action Modeling via Action-Conditioned Geometric Latent Prediction

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing world action models, which lack explicit geometric supervision and are prone to future information leakage through visual features. We propose ACG-WAM, which introduces an action-conditioned geometric Joint Embedding Predictive Architecture (JEPA) that applies geometric supervision prior to temporal mixing, targeting a frozen VGGT encoder. Furthermore, we pioneer an action-conditioned geometric latent prediction mechanism that eliminates information leakage and enables efficient deployment by removing auxiliary modules during inference. The approach is further enhanced by multi-camera shared visual embeddings for improved representation learning. Extensive evaluations demonstrate that ACG-WAM achieves a 93.46% success rate on the RoboTwin 2.0 benchmark and 85% on real-world robotic tasks, significantly outperforming the Motus baseline.
📝 Abstract
World action models jointly learn visual predictionand robot actions, providing a way to use observations ofscene evolution for policy learning. Their video and actionlosses, however, provide no explicit target for the geometricconsequences of a demonstrated action sequence. Moreover,visual features taken after temporal attention can contain futureobservations, making them unsuitable as the sole current visualinput to an auxiliary predictor. We introduce ACG-WAMand its auxiliary objective, the Action-Conditioned GeometricJoint-Embedding Predictive Architecture (ACG-JEPA), whichpredicts geometric features at several horizons from the currentobservation and intervening actions, using the future slot of afrozen VGGT encoding of each current and future image pairas the target. We apply this supervision from the head and wristcameras to a shared visual embedding before temporal mixing,and remove the teacher and auxiliary modules at inference.On 50 RoboTwin 2.0 tasks, ACG-WAM achieves 93.46%success in clean scenes, with the best randomized success(92.68%) and mean across both settings (93.07%) among thecompared methods; across three tasks on a real robot, itachieves 85.00% success and 91.67% partial completion score,exceeding Motus by 10.00 and 9.17 percentage points, respec-tively. Code:https://github.com/RoboOpus/ACG-WAM.Website:https://RoboOpus.github.io/ACG-WAM.
Problem

Research questions and friction points this paper is trying to address.

World Action Model
Geometric Prediction
Visual Representation
Policy Learning
Temporal Attention Leakage
Innovation

Methods, ideas, or system contributions that make the work stand out.

Action-Conditioned Geometric Prediction
Joint-Embedding Predictive Architecture
World Action Model
Temporal Attention
Robot Policy Learning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Jiangtao Liu
Jiangtao Liu
Intel Corporation
OptimizationSemiconductor Supply ChainManufacturingNetwork Flow ModelTransportation System
Z
Zishang Xiang
School of Automation, Beijing Institute of Technology
Y
Yage He
LimX Dynamics
L
Lingguo Cui
School of Automation, Beijing Institute of Technology
B
Baihai Zhang
School of Automation, Beijing Institute of Technology
R
Runqi Chai
School of Automation, Beijing Institute of Technology
S
Senchun Chai
School of Automation, Beijing Institute of Technology