🤖 AI Summary
This study addresses the limitations of existing world action models, which lack explicit geometric supervision and are prone to future information leakage through visual features. We propose ACG-WAM, which introduces an action-conditioned geometric Joint Embedding Predictive Architecture (JEPA) that applies geometric supervision prior to temporal mixing, targeting a frozen VGGT encoder. Furthermore, we pioneer an action-conditioned geometric latent prediction mechanism that eliminates information leakage and enables efficient deployment by removing auxiliary modules during inference. The approach is further enhanced by multi-camera shared visual embeddings for improved representation learning. Extensive evaluations demonstrate that ACG-WAM achieves a 93.46% success rate on the RoboTwin 2.0 benchmark and 85% on real-world robotic tasks, significantly outperforming the Motus baseline.
📝 Abstract
World action models jointly learn visual predictionand robot actions, providing a way to use observations ofscene evolution for policy learning. Their video and actionlosses, however, provide no explicit target for the geometricconsequences of a demonstrated action sequence. Moreover,visual features taken after temporal attention can contain futureobservations, making them unsuitable as the sole current visualinput to an auxiliary predictor. We introduce ACG-WAMand its auxiliary objective, the Action-Conditioned GeometricJoint-Embedding Predictive Architecture (ACG-JEPA), whichpredicts geometric features at several horizons from the currentobservation and intervening actions, using the future slot of afrozen VGGT encoding of each current and future image pairas the target. We apply this supervision from the head and wristcameras to a shared visual embedding before temporal mixing,and remove the teacher and auxiliary modules at inference.On 50 RoboTwin 2.0 tasks, ACG-WAM achieves 93.46%success in clean scenes, with the best randomized success(92.68%) and mean across both settings (93.07%) among thecompared methods; across three tasks on a real robot, itachieves 85.00% success and 91.67% partial completion score,exceeding Motus by 10.00 and 9.17 percentage points, respec-tively. Code:https://github.com/RoboOpus/ACG-WAM.Website:https://RoboOpus.github.io/ACG-WAM.