Topometric Autonomous Vehicle Localization by Combining Visual Embeddings and Feed-Forward 3D Models

πŸ“… 2026-08-06
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the challenge of simultaneously achieving robustness to appearance changes and high metric accuracy in visual autonomous vehicle localization. The authors propose a novel topological-metric fusion localization framework that integrates a feedforward 3D neural geometric model with probabilistic visual place recognition for the first time. The approach constructs a topology-metric map offline by associating poses with visual appearances and, during online operation, fuses visual place recognition outputs with 3D geometric trajectory estimates using particle filtering to achieve precise and perception-confusion-resistant localization. Leveraging low-dimensional visual embeddings and a sequential belief update mechanism, the system benefits from a modular design. Experiments demonstrate that the method significantly outperforms existing appearance-based localization approaches on three standard benchmarks, effectively mitigating localization failures caused by drastic environmental appearance changes.
πŸ“ Abstract
Effective Visual Localization (VL) requires a map of the environment that combines compactness for efficient scalability with robustness against visual appearance changes and metric precision. Through low-dimensional image embeddings, Visual Place Recognition (VPR) is able to successfully meet the first two requirements, but its low metric accuracy makes it less suitable than standard VL approaches based on local features or neural representations. This limitation can be overcome by integrating VPR with the accurate local trajectory estimates produced by feed-forward neural 3D geometry (FF3D) models. In this paper, we address sequential appearance-based localization through a topometric framework that iteratively combines probabilistic VPR with FF3D metric pose estimation in controlled image sets. Our approach proposes an automatic offline mapping tool that models the topometric pose-appearance interaction in the different parts of the scene. This map is later employed by an online particle filter that estimates the pose from odometry and belief over places for FF3D inference, successfully incorporating neural metric estimation into probabilistic appearance-based localization. We extensively evaluate the framework on three known benchmarks, demonstrating substantial improvements over existing appearance-based methods. The modularity of our approach allows the descriptor extractor and FF3D model to remain interchangeable, and a focused analysis further shows that sequential belief can mitigate severe failures under perceptual aliasing.
Problem

Research questions and friction points this paper is trying to address.

Visual Localization
Visual Place Recognition
Metric Accuracy
Topometric Mapping
Perceptual Aliasing
Innovation

Methods, ideas, or system contributions that make the work stand out.

topometric localization
visual place recognition
feed-forward 3D models
neural metric estimation
particle filter
πŸ”Ž Similar Papers
No similar papers found.