🤖 AI Summary
This study addresses the significant challenge of jointly reconstructing geometry and motion, as well as establishing cross-view correspondences, from multi-view and multi-temporal observations in dynamic scenes. To this end, it proposes ARROW, a feed-forward model whose core innovation lies in an order-invariant query mechanism that supports arbitrary input associations, thereby unifying 3D reconstruction and point tracking within a single framework. Furthermore, diverse training strategies are incorporated to substantially enhance generalization capabilities. The proposed approach establishes new state-of-the-art performance on both the WorldTrack and TAPVid-3D benchmarks, surpassing specialized multi-view trackers while maintaining strong competitiveness in 3D reconstruction tasks.
📝 Abstract
Dynamic scenes may be captured by a moving camera, multiple video streams, or images taken at different times. These observations reveal complementary aspects of scene geometry and motion, yet bringing them together requires establishing correspondence across viewpoints, capture times, and visibility changes. We introduce ARROW, a feed-forward model that unifies 3D reconstruction and 3D point tracking from arbitrary image sets. At its core is a novel order-invariant querying approach, which allows the association of queries with observations across arbitrary inputs. We show that exposing the model to more diverse sets of inputs during training results in improved task performance. Moreover, the resulting model is capable of generalization to a wider range of tasks including multi-view tracking. Trained with this strategy, ARROW establishes a new state of the art in 3D tracking on WorldTrack and TAPVid-3D and outperforms dedicated multi-view trackers on an adapted RGB-only MVTracker benchmark, while remaining competitive across 3D reconstruction tasks. Code and weights are publicly available.