How to Reduce Localization Ambiguity? Geometry-Semantic Constrained BEV Representation Learning for Satellite-Ground Localization

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the geometric and descriptor matching ambiguities inherent in bird’s-eye-view (BEV) feature placement for satellite-to-ground cross-view localization. To this end, we propose GeoSem-BEV, a method that introduces a dual-constraint mechanism integrating geometric and explicit semantic supervision. By jointly leveraging radial depth supervision, vertical height supervision, and cross-view semantic consistency, the proposed approach synergistically optimizes BEV representation learning, effectively resolving localization ambiguity while enhancing both positional accuracy and descriptor discriminability. Extensive experiments on benchmark datasets such as VIGOR demonstrate the efficacy of our method. Notably, under unknown orientation settings, GeoSem-BEV reduces the mean directional error by over 37%, yielding substantial improvements in cross-view localization performance.
📝 Abstract
Satellite-ground localization estimates the planar position and yaw orientation of a ground camera within a geo-referenced satellite image. Most recent methods map ground and satellite features into a shared bird's-eye-view (BEV) space and establish spatial correspondences. However, insufficient depth constraints can assign one ground feature to different distances along a viewing direction, creating geometric ambiguity in BEV feature placement. Similar appearances at different locations can also create descriptor matching ambiguity, while existing descriptor learning lacks explicit semantic supervision to distinguish them. We propose GeoSem-BEV, a geometry-semantic constrained BEV representation learning method. Radial depth supervision constrains distance assignment, and vertical height supervision constrains height aggregation. Shared explicit semantic supervision promotes consistent semantic predictions across views and helps distinguish locations with similar semantics. These constraints improve feature placement and descriptor discriminability, enhancing state-of-the-art BEV localization models. On VIGOR with unknown orientation, GeoSem-BEV reduces mean orientation error by 37.2% and 38.1% in the cross-area and same-area settings, respectively. The corresponding errors are reduced by 10.8% and 15.6% on DReSS-D. On KITTI-CVL, it reduces same-area mean orientation error by 26.8% under 10 degree orientation noise.
Problem

Research questions and friction points this paper is trying to address.

Satellite-Ground Localization
Localization Ambiguity
Bird's-Eye-View (BEV)
Geometric Ambiguity
Descriptor Matching Ambiguity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Satellite-Ground Localization
BEV Representation Learning
Geometry-Semantic Constraints
Depth Supervision
Descriptor Discriminability
🔎 Similar Papers
No similar papers found.