LoRetta: A Foundation Model and Extensive Dataset for Global-Scale Remote Sensing Dense Image Matching

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of dense matching in global multi-temporal remote sensing imagery, which suffers from large geometric offsets, partial overlaps, and non-matchable regions due to varying imaging conditions. The authors propose a two-stage “localization–registration” framework: first, a matchability-aware module identifies matchable regions and estimates an affine transformation; subsequently, dense residual correspondences are predicted within the aligned coordinate system under guided supervision. Key contributions include the creation of LEVIR-GM—the first global remote sensing matching benchmark with native matchability annotations—and a unified evaluation protocol, along with an efficient joint architecture that fuses multi-scale features. On LEVIR-GM, the method achieves an AUC of 83.3%, outperforming RoMa v2 by 1.6 points, improves PCK at ½-pixel by 6.5 and 8.2 points, reduces inference latency by 47.8%, and demonstrates strong generalization in cross-platform geolocation tasks.
📝 Abstract
Dense image matching establishes pixel-wise correspondences and underpins broad applications in computer vision and photogrammetry. However, extending dense matching to global-scale remote sensing remains challenging because image pairs may differ in acquisition time, season, viewpoint, spatial resolution, and land-cover state. The resulting large geometric offsets, partial overlap, and intrinsically unmatchable regions make direct dense correspondence prediction unreliable and inefficient. We thus reformulate dense matching as localization-and-registration: first localizing the matchable overlap and affine geometry, then refining dense residuals within the aligned frame. Based on this formulation, we propose LoRetta, a foundation model coupling matchability-aware affine localization with guided dense registration. We also introduce LEVIR-GM, a global-scale multi-temporal optical matching benchmark with dataset-native matchability labels (103K aligned, 827K augmented pairs, six continents, five years, 0.5-1024 m resolution). We further establish a unified evaluation protocol for sparse, semi-dense, and dense matchers. On LEVIR-GM, LoRetta achieves an area under the curve (AUC) of 83.3%, outperforming the strongest baseline RoMa v2 by 1.6 points, with larger percentage of correct keypoints (PCK) gains of 6.5 and 8.2 points at 1 and 2 pixels, while reducing inference latency by 47.8%. Astronaut-to-satellite and unmanned aerial vehicle (UAV)-to-satellite geolocalization experiments further demonstrate its transferability as a reusable geometric aligner.
Problem

Research questions and friction points this paper is trying to address.

dense image matching
remote sensing
global-scale
geometric correspondence
multi-temporal
Innovation

Methods, ideas, or system contributions that make the work stand out.

dense image matching
foundation model
affine localization
remote sensing
geometric registration
🔎 Similar Papers
No similar papers found.