Calibrating Retrieval Geometry: Reliability-Guided Training-Free Aggregation for Visual Place Recognition

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种无需训练的聚合方法TFA,通过可靠性引导来解决固定聚合在新环境中抑制有用区别的问题,提高了视觉地点识别的准确性。
📝 Abstract
Frozen visual foundation models provide transferable features for visual place recognition, but fixed aggregation can suppress useful distinctions in new environments. We introduce TFA, a reliability-guided, training-free aggregation method requiring neither place labels nor task-specific weight updates. Our key observation is that reproducible retrieval need not be discriminative: independent codebooks can consistently retrieve a few database hubs. TFA combines cross-codebook agreement, retrieval coverage, and spectral statistics to control residual assignment, spectral shaping, and global-feature fusion. Its spectral kernel exactly recovers original descriptor similarity at zero intervention. Database-only TFA fixes its rules before accessing queries; TFA-C64 uses 64 disjoint unlabeled target images to calibrate retrieval for subsequent queries. Across 20 ground protocols with a fixed DINOv2-B backbone and matched resolution, database-only TFA improves Recall@1 over AnyLoc by 17.39 percentage points on MSLS-val and 9.55 on SPED. C64 mitigates failures of database-only calibration in driving environments. Across eight aerial/cross-view protocols, TFA achieves the highest Recall@1 among compared training-free heads in 14 of 16 DINOv2/DINOv3 backbone-protocol combinations. In a separate native-system comparison, DINOv2-G-based TFA-C64 reaches 91.46% Recall@1 on Pitts30k and 76.29% on VPAIR, outperforming the displayed training-free comparators on all five benchmarks. These results show that reliability-guided aggregation can recover additional retrieval capability from frozen representations, providing a practical baseline for new environments with scarce place supervision.
Problem

Research questions and friction points this paper is trying to address.

Visual Place Recognition
Frozen Visual Foundation Models
Reliability-Guided Aggregation
Training-Free
Innovation

Methods, ideas, or system contributions that make the work stand out.

Training-Free Aggregation
Reliability-Guided
Cross-Codebook Agreement
Spectral Shaping
Global-Feature Fusion
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xin Li
Key Laboratory of Spectral Imaging Technology CAS, Xi’an Institute of Optics and Precision Mechanics, Chinese Academy of Sciences, Xi’an, 710119, China; University of Chinese Academy of Sciences, Beijing, 100049, China
Z
Zhimin Mao
Key Laboratory of Spectral Imaging Technology CAS, Xi’an Institute of Optics and Precision Mechanics, Chinese Academy of Sciences, Xi’an, 710119, China; University of Chinese Academy of Sciences, Beijing, 100049, China
Shang Wang
Shang Wang
Stevens Institute of Technology
BiophotonicsFunctional Optical ImagingDevelopmentReproductive BiologyBiomechanics
S
Siyuan Duan
Key Laboratory of Spectral Imaging Technology CAS, Xi’an Institute of Optics and Precision Mechanics, Chinese Academy of Sciences, Xi’an, 710119, China; University of Chinese Academy of Sciences, Beijing, 100049, China
Geng Zhang
Geng Zhang
Northwestern Polytechnical University
IoT-enabled manufacturingSmart manufacturing system