Score
Designs, builds, and evaluates systems that determine whether a visual observation (image or video frame) corresponds to a previously seen physical place by producing place matches, labels, or retrieval rankings. Work covers creating and analyzing visual representations, feature descriptors, and matching/verification pipelines that are robust to viewpoint, illumination, seasonal, and dynamic scene changes for tasks such as loop-closure detection and image-based localization.
This work systematically re-examines the necessity of match re-ranking in visual place recognition (VPR). We observe that blind re-ranking degrades performance in modern high-accuracy retrieval systems operating on saturated datasets. To address this, we propose a “matching-as-verification” paradigm: using the number of inlier matches as a confidence metric, we design an adaptive, conditionally triggered re-ranking decision mechanism—replacing rigid, fixed pipelines. Our method integrates SuperPoint+SuperGlue for feature extraction, RANSAC-based geometric verification, and inlier count modeling. Evaluated on Nordland and Oxford RobotCar benchmarks, it reduces erroneous re-ranking by over 40%, significantly enhancing robustness while preserving high recall. This study is the first to reveal the potential harm of indiscriminate re-ranking in state-of-the-art VPR systems and establishes an interpretable, generalizable criterion for re-ranking decisions.
This work reframes visual place recognition (VPR) as an image-pair retrieval task to better support downstream applications such as scene registration, SLAM, and structure-from-motion. The study systematically evaluates prominent VPR methods—including NetVLAD, CosPlace, EigenPlaces, MixVPR, AnyLoc, SALAD, and MegaLoc—on three cross-domain 3D datasets: Tanks and Temples, ScanNet-GS, and KITTI. For the first time, VPR is explicitly modeled as a front-end for image-pair retrieval, revealing the domain dependency of existing approaches under challenges like perceptual aliasing and sequence incompleteness. Experimental results demonstrate that modern global descriptors can serve as plug-and-play, efficient retrieval modules, offering practical guidance for robust 3D mapping and registration pipelines.
This study addresses the insufficient robustness and accuracy of local feature matching in overlapping regions of satellite imagery. To this end, the authors construct a manually curated satellite image dataset annotated with GPS coordinates and conduct a systematic evaluation of SIFT and ORB algorithms across the entire matching pipeline—including keypoint detection, descriptor extraction, feature matching, and RANSAC-based geometric verification. Using the inlier ratio as the primary metric for matching quality, the work quantitatively analyzes the impact of keypoint quantity on matching performance. The results reveal a nonlinear relationship between the number of detected keypoints and the inlier ratio, offering empirical evidence and theoretical guidance for algorithm selection and parameter tuning in remote sensing image matching tasks.
This work addresses the failure of conventional SIFT-based image registration in scenes dominated by strong linear structures, where local features become ambiguous and poorly discriminative. To overcome this limitation, the authors propose a novel approach that, for the first time, transfers SIFT descriptors into Hough space for matching. By leveraging the Hough transform, linear structures are mapped to prominent peaks, thereby restoring the distinctiveness of features. The resulting Hough-space feature matching framework significantly outperforms standard SIFT in highly structured environments while maintaining comparable registration accuracy in general scenes. This method effectively mitigates the performance degradation of SIFT in structured settings, offering a robust solution for reliable image alignment across diverse scenarios.
This paper addresses topological localization of mobile robots in office environments using only monocular optical camera images. To tackle challenges including perceptual ambiguity, illumination variations, and sensor noise, we propose a robust and efficient visual localization framework. Our method systematically and quantitatively compares multiple visual descriptors (color histograms, SIFT, ASIFT, RGB-SIFT, BoW), similarity metrics, and classifiers—evaluated on the ImageCLEF dataset. Using standard evaluation metrics and visualization-based analysis, we identify the optimal descriptor–metric–classifier configuration. Experimental results demonstrate significant improvements in localization accuracy and generalization across varying illumination conditions and long-trajectory scenarios. The framework reliably maps novel image sequences to topological nodes without GPS or temporal constraints. It provides a reproducible, scalable solution for vision-only topological localization in real-world office settings.
This study investigates how synthesizing novel viewpoints can enhance visual place recognition (VPR) performance, particularly in cooperative navigation scenarios involving ground and aerial robots. The authors systematically evaluate the impact of various viewpoint synthesis strategies by integrating seven representative image similarity methods across five public VPR datasets. For the first time, they quantitatively analyze the combined effects of the number of synthesized views, the magnitude of viewpoint variation, and image type on VPR accuracy. Experimental results demonstrate that even a small number of synthesized views significantly improves recognition performance. Moreover, as the volume of synthesized data increases, the number of views and image type become dominant factors, surpassing the influence of viewpoint variation magnitude.
This work addresses the vulnerability of visual place recognition (VPR) in dynamic environments, where reliance on fixed matching thresholds often leads to erroneous loop closures due to the absence of ground-truth labels, thereby compromising the reliability of localization and mapping. To mitigate this issue, the authors propose a general, model-agnostic post-hoc verification framework that, for the first time, integrates vision-language models (VLMs) into VPR auditing. By leveraging cross-modal joint reasoning, the framework performs instance-level match verification between query and candidate images without requiring dataset-specific confidence thresholds, environment priors, or calibrated scores. Evaluated across six benchmark datasets, the method improves recall@1 by an average of 13.6%, reduces the false acceptance rate to 12%, and maintains precision above 95% with coverage exceeding 75%, significantly enhancing the robustness and safety of VPR systems.
This study addresses the unclear incremental value of street view imagery over existing data sources for urban perception using vision-language models (VLMs), as well as the accuracy-centric bias in conventional evaluations. By employing multi-source data fusion and spatial distance sensitivity analysis, this work compares image-based predictions with established urban datasets across seven attributes, systematically investigating how image visibility and data coverage influence predictive performance. The findings reveal that the utility of imagery depends critically on attribute visibility and the coverage rate of existing data; specifically, images demonstrate advantages only for certain attributes such as building typology, while non-image data outperform them in other scenarios. Furthermore, this research contributes OpenFACADES, a newly annotated dataset released to support future investigations in VLM-driven urban analytics.
VideoReloc通过适应性剪辑和假设优先注册等方法,解决了长期室内视频重定位问题,提高了在光照和家具变化下的定位成功率。
This work addresses the limitations of existing vision datasets, which predominantly rely on low-resolution, provenance-uncertain JPEG images and lack high-fidelity visual content and spatial context from real-world environments. To bridge this gap, the authors introduce a large-scale, high-fidelity scene dataset comprising 67,574 images captured across 810 real-world physical locations spanning 260 indoor, outdoor, and natural scene categories. Using a Canon EOS R5 camera, images were acquired at 5-degree horizontal intervals with multiple elevation angles, yielding synchronized 14-bit CR3 RAW and corresponding JPEG pairs alongside complete EXIF metadata. This dataset establishes a new ecologically valid and quality-controlled benchmark for research on viewpoint-dependent recognition in humans and models, real-world scene understanding, statistical analysis of natural images, and full-field-of-view vision experiments.