FAST: Flow Any Scene Transformer

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited scalability of vision foundation models in precise correspondence matching by proposing a parameter-free rewiring strategy to construct a ViT-based matcher. The method repurposes the query-key projections from single-view pretrained models as cross-view matching priors, achieving architectural innovation by converting self-attention into cross-attention, and integrates large-scale dense displacement estimation for efficient matching. Experimental results demonstrate that the proposed approach achieves state-of-the-art performance across multiple benchmarks. Furthermore, matching accuracy improves significantly with increasing backbone capacity and data volume, validating its superior scalability for precise correspondence tasks.
📝 Abstract
Scaling has become a primary driver of progress in language and vision foundation models, yet its role in precise correspondence matching remains underexplored. In this work, we present Flow Any Scene Transformer (FAST), a scalable correspondence model driven by two key insights. First, we reveal that the query-key projections inside single-view vision foundation models encode a coarse yet reusable prior for cross-view matching. Second, reusing these pretrained projections in cross-attention form yields a highly effective initialization for a ViT-based matcher built from a single-view encoder. Guided by these insights, we build FAST upon a vanilla single-view foundation model, utilizing a zero-parameter rewiring strategy to convert selected self-attention layers into cross-attention for cross-view interaction. This design allows ViT-based matchers to scale with advances in single-view foundation models, bypassing the need for a dedicated pair-centric pretraining stage. To fully unlock the scaling potential of this formulation, we assemble a 6-million-pair training corpus for general-purpose dense 2D displacement estimation across diverse co-visible image pairs. Extensive experiments demonstrate that FAST achieves state-of-the-art performance across a wide range of benchmarks, while scaling favorably with both backbone size and training data.
Problem

Research questions and friction points this paper is trying to address.

correspondence matching
scaling
dense 2D displacement estimation
vision foundation models
cross-view matching
Innovation

Methods, ideas, or system contributions that make the work stand out.

Scalable Correspondence Matching
Zero-parameter Rewiring
Vision Foundation Model
Cross-attention
Dense 2D Displacement Estimation
🔎 Similar Papers
No similar papers found.