World Motion Models: Flexible Sequence Modeling of SE(3) Trajectories

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of generative priors for dynamic 3D worlds and the challenge of unifying multi-entity motion modeling by proposing a general-purpose 4D representation framework that utilizes sparse SE(3) trajectories as minimal expressive primitives. The proposed method reformulates the joint distribution of diverse entities—including objects, humans, and robots—into a flexible sequence modeling problem. Furthermore, it incorporates flow matching techniques with token-wise noise scheduling and context mechanisms to enable non-autoregressive any-to-any conditional control. Experimental results demonstrate that the model exhibits superior cross-modal generalization across six 3D vision and robotics tasks, achieving state-of-the-art performance in scenarios such as motion prediction and inverse dynamics.
📝 Abstract
Equipping artificial agents with spatial intelligence requires a comprehensive generative prior over the dynamic 3D world. We propose World Motion Models (WMMs) that capture "what was, is, and will be where across time" via sparse SE(3) pose trajectories. WMMs are built on the observation that elements of dynamic scenes can be well approximated by a set of rigid SE(3) trajectories, a minimal yet expressive primitive for 4D modeling. This representation unifies articulated objects, human bodies, hand-object interactions, piecewise-rigid scene dynamics, camera motion, and even robot states and actions into a single shared space. Given this representation, we cast the joint distribution of these entities as a flexible sequence modeling problem, utilizing flow-matching with per-token noise levels. Coupled with a context token mechanism for non-sequential conditioning, this formulation supports any-to-any marginal conditioning across an arbitrary number of entities and time steps. Tasks such as future prediction, motion infilling, model-predictive control, inverse kinematics, cross-embodiment retargeting, and policy learning all reduce to the application of different masks over the same network. Experiments on 6 diverse applications of 3D vision and robotics demonstrate the versatility and flexibility of WMMs with strong performance.
Problem

Research questions and friction points this paper is trying to address.

spatial intelligence
SE(3) trajectories
4D modeling
sequence modeling
generative prior
Innovation

Methods, ideas, or system contributions that make the work stand out.

World Motion Models
SE(3) trajectories
Flow matching
Flexible sequence modeling
Any-to-any conditioning