LAWM-3D: Learning 3D-Aware Latent Actions from Human Videos for Generalizable Robot World Models

📅 2026-08-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing latent action models are constrained by single-view videos and 2D representations, hindering the development of general-purpose robot world models with 3D awareness, while also facing challenges such as high annotation costs and heterogeneous action spaces across platforms. This work proposes a two-stage framework leveraging self-supervised learning from multi-view human videos: first pretraining on large-scale multi-view video data, then fine-tuning for robotic tasks. The core innovations include a triply coupled design—view-invariant unified action tokenization, geometric alignment constraints derived from pretrained 3D foundation models, and a non-injective RGB-D joint reconstruction objective—which collectively suppress appearance leakage and enhance geometric motion modeling. The approach substantially improves the generative quality, physical consistency, and cross-task generalization of world models, achieving state-of-the-art performance across multiple metrics.
📝 Abstract
World models enable agents to perform forward rollout and planning without real-world interaction. However, their application in open-world embodied intelligence remains limited by the high cost of action annotations and the heterogeneity of action spaces across platforms. Recently, latent action models (LAMs) have alleviated this bottleneck by learning action representations directly from unlabeled human videos in a self-supervised manner. Nevertheless, most existing LAMs rely on single-view inputs and operate primarily in 2D pixel space, raising a fundamental question: can simply incorporating multi-view videos into LAM training endow the learned latent actions with 3D-aware perception? Our study shows that the answer is negative. The primary reasons lie in future-frame appearance leakage as well as inter-camera appearance discrepancies and viewpoint variations. To address these issues, we propose LAWM-3D, which introduces three tightly coupled key designs: (1) a multi-view invariant unified action tokenization scheme for learning 3D-aware latent actions; (2) a geometric alignment constraint that anchors intermediate encoder features to a pretrained 3D foundation model, thereby explicitly providing cross-view geometric correspondences; and (3) a non-injective RGB-D joint reconstruction objective that prevents shortcut learning from future-frame appearance information, forcing the LAM to focus supervision on motion cues with geometric significance. Importantly, these components are not simply stacked but are tightly coupled through a unified motivation. Built upon a two-stage paradigm of large-scale human video pretraining followed by robot fine-tuning, extensive experiments demonstrate that the proposed 3D-aware latent actions significantly improve world model performance, achieving SOTA results in generation quality, physical consistency, and generalization ability.
Problem

Research questions and friction points this paper is trying to address.

latent action models
3D-aware perception
multi-view videos
world models
embodied intelligence
Innovation

Methods, ideas, or system contributions that make the work stand out.

3D-aware latent actions
latent action models
multi-view video
geometric alignment
world models
🔎 Similar Papers
No similar papers found.