3D Point Tracking with State Space Models

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the dual challenges of low accuracy and high computational overhead in absolute metric 3D point tracking under monocular, pose-free conditions. To this end, it proposes a feed-forward framework that integrates dense optical flow, monocular metric depth estimation, and DINOv3 features. Central to this approach is the introduction of the Mamba-3 state space model, which replaces conventional Transformer architectures to optimize pixel-ray depth residuals. This architectural substitution reduces memory complexity from linear to constant, enabling efficient inference on a single GPU. Evaluated on the TAPVid-3D benchmark, the proposed method achieves state-of-the-art absolute metric accuracy with an mAJ of 0.256, substantially outperforming existing feed-forward approaches while effectively balancing high precision with low resource consumption.
📝 Abstract
Tracking any point of a dynamic scene in metric 3D - in absolute meters, not up to an unknown scale - underpins 3D and 4D reconstruction, robot navigation, and autonomous driving, where decisions are made in meters, not pixels. Our objective is a 3D point tracker accurate in those absolute terms and operating within a single commodity GPU, pose-free, monocular budget. Our method rests on one observation: once a point's 2D image trajectory is fixed, the quantity that governs its metric accuracy is the depth along its pixel ray. Rather than learning tracking end-to-end, we therefore compose two frozen front-ends - dense optical flow for 2D correspondence and a monocular metric-depth network for the third dimension - and learn only the residual they cannot supply: that depth, refined by a compact state space model (Mamba-3) conditioned on appearance features (DINOv3). A state space model rather than the transformers the strongest 3D trackers adopt is what makes a single-GPU budget attainable: it summarises a track in a fixed-size recurrent state whose memory cost is constant in the number of frames, whereas attention requires a key-value cache that grows linearly with them. On the TAPVid-3D minival benchmark our best configuration attains the highest absolute metric accuracy among methods evaluated under identical conditions (mean metric Average Jaccard, 0.256), exceeding strong feed-forward trackers, while a companion analysis, reproduced with each competitor's own evaluator, explains why several published trackers lose most of their accuracy under this budget.
Problem

Research questions and friction points this paper is trying to address.

3D point tracking
metric depth
monocular
state space models
single GPU
Innovation

Methods, ideas, or system contributions that make the work stand out.

State Space Models
3D Point Tracking
Metric Depth
Monocular
Mamba
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
M
Masahiro Ogawa
Department of Precision Engineering, Graduate School of Engineering, The University of Tokyo, Kashiwa, Chiba, Japan
Qi An
Qi An
The University of Tokyo
Robotics
A
Atsushi Yamashita
Department of Human and Engineered Environmental Studies, Graduate School of Frontier Sciences, The University of Tokyo, Kashiwa, Chiba, Japan