🤖 AI Summary
This study addresses the challenge that point tracking models struggle to simultaneously handle long-term sparse and short-term dense tracking. To this end, it formulates video modeling as persistent 3D scene trajectories in world coordinates, thereby decoupling model complexity from video duration. Methodologically, the proposed approach replaces 4D correlation volumes with 3D point cloud representations and introduces sliding-window voxel deduplication, endpoint-trajectory decomposition refinement, and 3D WAFT feature sampling, integrated with dynamic-static classification to achieve efficient dense tracking. As the first 3D tracker capable of tracking all visible points across over a thousand frames within 40GB of VRAM, this method yields an improvement exceeding 20% on the Average Point Density (APD) metric for short clips.
📝 Abstract
Existing point tracking models face a fundamental tradeoff: they can either track a sparse set of query points over long horizons, or track all points across only short clips. We introduce TrackEverything, a 3D point tracker that breaks this trade-off by representing videos as persistent 3D scene tracks in world coordinates. Grounded in the insight that videos are 2D projections of an underlying 3D world, TrackEverything decouples model complexity from video duration, allowing it to scale with unique physical scene geometry instead. Our approach introduces three key innovations. First, we employ a voxelization-based de-duplication mechanism at sliding-window boundaries to merge co-located tracks, preventing repeated observations of the same surface from redundantly accumulating. Second, we decompose tracking into an endpoint refiner that predicts each point's destination and static-versus-dynamic classification, followed by a lightweight trajectory refiner that decodes dense trajectories exclusively for dynamic points. Third, we propose 3D WAFT, replacing memory-prohibitive 4D correlation volumes with efficient feature sampling in the scene cloud. To the best of our knowledge, TrackEverything is the first 3D tracker capable of tracking all visible points across videos exceeding 1000 frames within 40 GB of GPU memory. On TAPVid-3D, TrackEverything outperforms all open-source all-frame dense 3D trackers by more than 20% APD on short clips, while remaining competitive with state-of-the-art sparse trackers on long sequences, despite tracking far more points.