Focus What Matters: Matchability-Based Reweighting for Local Feature Matching

📅 2025-05-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing semi-dense matching methods apply uniform weighting to all pixels during attention-based feature extraction, making them susceptible to noise from redundant and irrelevant regions and yielding poorly discriminative attention weights. To address this, we propose a matchability-aware dual-path dynamic reweighting mechanism: (i) injecting learnable matchability bias into attention logits, and (ii) performing matchability-driven post-attention adaptive scaling of value features. This mechanism employs a lightweight binary classification head to estimate pixel-wise matchability in real time, enabling fine-grained, semantics-aware attention modulation. Fully embedded within the Transformer architecture, our method requires no additional supervision or pretraining. Evaluated on HPatches, ETH3D, and SUN3D—three major benchmarks—it consistently outperforms state-of-the-art methods, achieving significant improvements in repeatability, matching accuracy, and robustness to viewpoint and illumination variations.

Technology Category

Computer Vision: Multi-modal VisionMachine Learning: Feature Construction/ReformulationIntelligent Robots: Multimodal Perception & Sensor Fusion

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and ranking
📝 Abstract
Since the rise of Transformers, many semi-dense matching methods have adopted attention mechanisms to extract feature descriptors. However, the attention weights, which capture dependencies between pixels or keypoints, are often learned from scratch. This approach can introduce redundancy and noisy interactions from irrelevant regions, as it treats all pixels or keypoints equally. Drawing inspiration from keypoint selection processes, we propose to first classify all pixels into two categories: matchable and non-matchable. Matchable pixels are expected to receive higher attention weights, while non-matchable ones are down-weighted. In this work, we propose a novel attention reweighting mechanism that simultaneously incorporates a learnable bias term into the attention logits and applies a matchability-informed rescaling to the input value features. The bias term, injected prior to the softmax operation, selectively adjusts attention scores based on the confidence of query-key interactions. Concurrently, the feature rescaling acts post-attention by modulating the influence of each value vector in the final output. This dual design allows the attention mechanism to dynamically adjust both its internal weighting scheme and the magnitude of its output representations. Extensive experiments conducted on three benchmark datasets validate the effectiveness of our method, consistently outperforming existing state-of-the-art approaches.
Problem

Research questions and friction points this paper is trying to address.

Classify pixels into matchable and non-matchable categories
Propose attention reweighting mechanism with learnable bias
Improve local feature matching by reducing noisy interactions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Classifies pixels into matchable and non-matchable categories
Introduces learnable bias term for selective attention adjustment
Applies matchability-informed rescaling to input value features
💼 Related Jobs
No related jobs found.
D
Dongyue Li