Geometry-Aligned Semantic Matching for Cross-Modal Planar Image Registration

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of semantic-geometric inconsistency and the absence of global guidance in cross-modal matching by proposing the CDPM framework. This work is the first to introduce geometric consistency constraints into DINOv3 feature adaptation, constructing a DINO-centric multi-scale feature pyramid through geometry-aware patch pair mining. By integrating a lightweight CNN for local structure refinement, the method establishes a semantics-led, detail-assisted matching architecture that overcomes the conventional stability-accuracy trade-off. Evaluated on VIS-IR datasets, CDPM achieves an AUC improvement exceeding 10 percentage points and reduces the mean angular error to 2.78 pixels. Furthermore, it outperforms RoMa v2 while reducing computational cost by 45.6%.
📝 Abstract
Cross-modal image matching establishes stable and accurate geometric correspondences across modalities for planar registration. Existing semantic representations provide cross-modal consistency, but semantic similarity does not necessarily imply geometric correspondence. Meanwhile, fine-grained CNN features provide accurate local details but lack global cross-modal semantic guidance for stable refinement. To address these issues, we propose CDPM, which first establishes geometrically consistent semantic representations and then preserves their dominant role in correspondence estimation during fine-grained localization. Specifically, we progressively adapt DINOv3 using geometrically consistent cross-modal patch pairs, enabling feature similarity to better reflect true cross-modal spatial correspondences. We then construct a DINO-Centric Feature Pyramid, where multi-scale DINO representations maintain stable cross-modal correspondences, while a lightweight CNN branch provides auxiliary structural details for precise local refinement. Extensive experiments on three cross-modal datasets demonstrate the superior performance of CDPM. On VIS-IR, compared with the dense matcher RoMa, CDPM improves AUC@3/5/10/20 by 7.36, 13.40, 13.75, and 10.42 percentage points, respectively, and reduces mACE from 5.83 to 2.78 pixels. It also outperforms RoMa v2 across all metrics while requiring 45.6% fewer FLOPs. The online demo and dataset are available, and the code will be released on our project page at https://warren-wzw.github.io/CDPM/.
Problem

Research questions and friction points this paper is trying to address.

Cross-modal image registration
Semantic matching
Geometric correspondence
Feature alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-Modal Image Registration
Geometric-Aligned Semantic Matching
DINOv3 Adaptation
Feature Pyramid
Planar Image Registration
💼 Related Jobs
No related jobs found.