🤖 AI Summary
This study addresses the limited scalability of vision foundation models in precise correspondence matching by proposing a parameter-free rewiring strategy to construct a ViT-based matcher. The method repurposes the query-key projections from single-view pretrained models as cross-view matching priors, achieving architectural innovation by converting self-attention into cross-attention, and integrates large-scale dense displacement estimation for efficient matching. Experimental results demonstrate that the proposed approach achieves state-of-the-art performance across multiple benchmarks. Furthermore, matching accuracy improves significantly with increasing backbone capacity and data volume, validating its superior scalability for precise correspondence tasks.
📝 Abstract
Scaling has become a primary driver of progress in language and vision foundation models, yet its role in precise correspondence matching remains underexplored. In this work, we present Flow Any Scene Transformer (FAST), a scalable correspondence model driven by two key insights. First, we reveal that the query-key projections inside single-view vision foundation models encode a coarse yet reusable prior for cross-view matching. Second, reusing these pretrained projections in cross-attention form yields a highly effective initialization for a ViT-based matcher built from a single-view encoder. Guided by these insights, we build FAST upon a vanilla single-view foundation model, utilizing a zero-parameter rewiring strategy to convert selected self-attention layers into cross-attention for cross-view interaction. This design allows ViT-based matchers to scale with advances in single-view foundation models, bypassing the need for a dedicated pair-centric pretraining stage. To fully unlock the scaling potential of this formulation, we assemble a 6-million-pair training corpus for general-purpose dense 2D displacement estimation across diverse co-visible image pairs. Extensive experiments demonstrate that FAST achieves state-of-the-art performance across a wide range of benchmarks, while scaling favorably with both backbone size and training data.