🤖 AI Summary
This study addresses the challenging problem of absolute metric localization and heading estimation for UAVs using only a single oblique image against satellite maps. We propose SatFix, a framework built upon the VGGT-Ω architecture that pioneers the reformulation of cross-view retrieval as continuous coordinate and heading regression. By employing satellite grid features as queries to aggregate visual evidence, SatFix directly outputs 3-DoF poses through a lightweight head network, enabling feed-forward localization without 3D maps or auxiliary sensors while supporting unified single- and multi-view modeling. Furthermore, we construct University-Metric, a large-scale evaluation benchmark. Experimental results demonstrate that the single-view model achieves a 52% localization rate within 50 meters with inference under 0.1 seconds, while the multi-view extension reduces positional error by 34% and refines heading error to 8.73°.
📝 Abstract
We study absolute metric UAV localization within a provided geo-referenced satellite region, recovering continuous map position and viewing heading from a single oblique image or a short multi-view clip. Existing cross-view geo-localization methods retrieve the most similar satellite tile from a gallery and report Recall@K, but retrieval depends on gallery sampling, provides no heading estimate, and returns a tile index rather than a continuous coordinate. We propose SatFix, a feed-forward UAV--satellite localization framework built on VGGT-$Ω$. Satellite-grid features act as queries that aggregate UAV visual evidence, and two lightweight heads regress a 3-DoF pose in the satellite-map frame: continuous 2D position and heading. SatFix requires no explicit 3D map, rendered bird's-eye image, auxiliary sensor, or test-time pose alignment. A single model supports both single- and multi-view inputs, with trajectory constraints used during multi-view training. For metric evaluation, we introduce University-Metric, where satellite imagery is re-collected over a region up to 10.7$\times$ longer on a side (about 114$\times$ the ground area) than the original University-1652 tiles, with continuous position and heading labels for the original UAV tours. With one UAV view, SatFix localizes 52.08% of test frames within 50 m and 17.34% within 10 m, with median position and heading errors of 45.66 m and $20.81^\circ$, respectively. Inference takes under 0.1 s per single-view query on an NVIDIA RTX 4090. With nine UAV views, the median position error falls to 21.96 m and the median heading error to $8.73^\circ$. Compared with a fine-tuned VGGT-$Ω$ baseline, SatFix reduces median position error by 34.0% and nine-view median heading error from $25.43^\circ$ to $8.73^\circ$.