🤖 AI Summary
This study addresses the dual challenges of low accuracy and high computational overhead in absolute metric 3D point tracking under monocular, pose-free conditions. To this end, it proposes a feed-forward framework that integrates dense optical flow, monocular metric depth estimation, and DINOv3 features. Central to this approach is the introduction of the Mamba-3 state space model, which replaces conventional Transformer architectures to optimize pixel-ray depth residuals. This architectural substitution reduces memory complexity from linear to constant, enabling efficient inference on a single GPU. Evaluated on the TAPVid-3D benchmark, the proposed method achieves state-of-the-art absolute metric accuracy with an mAJ of 0.256, substantially outperforming existing feed-forward approaches while effectively balancing high precision with low resource consumption.
📝 Abstract
Tracking any point of a dynamic scene in metric 3D - in absolute meters, not up to an unknown scale - underpins 3D and 4D reconstruction, robot navigation, and autonomous driving, where decisions are made in meters, not pixels. Our objective is a 3D point tracker accurate in those absolute terms and operating within a single commodity GPU, pose-free, monocular budget. Our method rests on one observation: once a point's 2D image trajectory is fixed, the quantity that governs its metric accuracy is the depth along its pixel ray. Rather than learning tracking end-to-end, we therefore compose two frozen front-ends - dense optical flow for 2D correspondence and a monocular metric-depth network for the third dimension - and learn only the residual they cannot supply: that depth, refined by a compact state space model (Mamba-3) conditioned on appearance features (DINOv3). A state space model rather than the transformers the strongest 3D trackers adopt is what makes a single-GPU budget attainable: it summarises a track in a fixed-size recurrent state whose memory cost is constant in the number of frames, whereas attention requires a key-value cache that grows linearly with them. On the TAPVid-3D minival benchmark our best configuration attains the highest absolute metric accuracy among methods evaluated under identical conditions (mean metric Average Jaccard, 0.256), exceeding strong feed-forward trackers, while a companion analysis, reproduced with each competitor's own evaluator, explains why several published trackers lose most of their accuracy under this budget.