MoSE3: Learning World-Space SE(3) at Every Pixel

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of rotational and rigid grouping information in dense 3D point tracking from monocular videos by proposing the first feedforward, per-pixel 6-DoF rigid motion prediction model. The approach jointly learns 3D trajectories and rigidity embeddings, integrating SE(3) manifold regression with differentiable soft-clustering optimization to recover per-pixel SE(3) motions in world coordinates end-to-end. Furthermore, a large-scale synthetic dataset, Art-Kubric, is constructed to bridge the annotation gap. Experimental results demonstrate that the proposed model achieves state-of-the-art SE(3) estimation accuracy and 3D tracking performance on both rigid and articulated object benchmarks, while exhibiting strong generalization capabilities to real-world scenarios.
📝 Abstract
Dense 3D point tracking has been a prominent paradigm for modeling motion in dynamic scenes, but a point track is just a 3-DoF translation curve per pixel: it captures where pixels go, not the rotation of the underlying part, nor which pixels move together as one body. We propose MoSE3, the first feed-forward model that predicts dense SE(3) motion from monocular RGB video, producing full 6-DoF rigid transforms at every pixel in world space. Per-pixel SE(3) motion offers a richer view of how a scene moves: rotation, translation, and grouping all at once. Directly predicting SE(3) is challenging: rotations lie on a curved manifold that is ill-suited to Euclidean regression, and annotations for SE(3) are particularly difficult to acquire. To address these challenges, MoSE3 predicts per-pixel SE(3) through two jointly learned intermediates, 3D point tracks and rigidity embeddings, and recovers SE(3) by differentiably fitting transforms within each soft rigid cluster, enabling end-to-end prediction and supervision. To close the data gap, we introduce Art-Kubric, a large-scale synthetic dataset with dense SE(3) and rigidity labels for articulated objects with rich physical interactions. MoSE3 achieves state-of-the-art SE(3) estimation at pixel, part, and object levels on both rigid and articulated benchmarks, and state-of-the-art average 3D point tracking accuracy across three datasets, while showing strong generalization to real-world videos despite being trained solely on synthetic motion data.
Problem

Research questions and friction points this paper is trying to address.

dense 3D point tracking
per-pixel SE(3) motion
monocular RGB video
rigid body motion
articulated objects
Innovation

Methods, ideas, or system contributions that make the work stand out.

SE(3) motion estimation
dense 3D point tracking
rigidity embeddings
differentiable transform fitting
Art-Kubric dataset
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.