🤖 AI Summary
This study addresses the excessive computational and memory overhead caused by full token matrices in semi-dense matching. We propose an efficient image matching framework based on candidate path routing. Specifically, a lightweight transport path router combined with block-level representation ranking is designed to filter candidate paths, and a sparse global dual-Softmax mechanism is introduced for efficient matching. Furthermore, structural reparameterization and shared-parameter fine-tuning heads are employed to optimize deployment performance. Experimental results demonstrate that the proposed method achieves a 1.67× speedup over SuperPoint+LightGlue while requiring only 0.44 GiB of GPU memory. It supports real-time inference at 6K resolution on a single GPU, effectively unifying accuracy, efficiency, and scalability.
📝 Abstract
Despite recent advances in accuracy and efficiency, coarse matching remains an indispensable yet costly stage in existing semi-dense matchers due to dense token-level matching. We present UltraMatch, an ultra-efficient and scalable semi-dense matching framework that bypasses the quadratic computation and memory cost of dense token-level matching by routing only a small fraction of candidate matching paths. At its core, a lightweight Transport Path Router operates on coarse block representations to rank candidate target blocks for each source block and retain only a small set, restricting subsequent token-level matching to the selected paths and avoiding the construction of the full token-to-token matching matrix. We further design a sparse global Dual-Softmax that performs matching only over the routed block candidates while retaining global competition across the sparse matching space. Beyond matching acceleration, UltraMatch employs deployment-oriented structural reparameterization for feature extraction and a tiny fine matching head with shared parameters, further reducing inference cost and memory consumption. UltraMatch achieves competitive accuracy among semi-dense matchers, while running 1.67$\times$ faster than SuperPoint+LightGlue with only 0.44 GiB peak inference memory. Its scalability enables inference at up to 6K resolution on a single RTX 3090, whereas existing semi-dense matchers run out of memory before reaching 2K. Our routing strategy is also transferable, delivering about 2$\times$ end-to-end speedup in EDM and ELoFTR without accuracy loss. The project repository is available at https://github.com/JiajunLe/UltraMatch.