Are All Tokens Necessary for Visual Place Recognition? An Empirical Study of Token Reduction for Efficient Inference

📅 2026-07-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of deploying vision Transformer-based models for visual place recognition (VPR) in real-time or resource-constrained settings due to their high computational overhead. To this end, the study establishes the first systematic benchmark for token reduction in VPR, comprehensively evaluating token pruning, merging, and hybrid strategies across multiple state-of-the-art architectures and real-world scenarios in terms of the trade-off between efficiency and accuracy. Experimental results demonstrate that the explored methods can reduce computational cost by up to 29% and increase throughput by up to 44%, with less than 1% degradation in recognition accuracy. These findings reveal significant token redundancy inherent in VPR tasks and provide practical guidance for efficient model deployment in real-world applications.
📝 Abstract
Recent visual place recognition (VPR) methods based on vision transformers, particularly foundation models, have achieved remarkable recognition performance. However, these models process all visual tokens throughout the entire network, resulting in substantial computational overhead, which hinders their deployment in real-time and resource-constrained scenarios. A natural question thus arises: are all visual tokens necessary for VPR? To answer this question, we present the first systematic benchmark of token reduction for efficient visual place recognition. Our benchmark comprehensively evaluates representative token pruning, token merging, and hybrid pruning-merging methods across multiple state-of-the-art VPR models and diverse benchmark datasets covering urban, suburban, and natural environments. We further investigate token reduction from multiple perspectives, including recognition performance under different reduction configurations, computational complexity, inference speed, qualitative visualization, and deployment efficiency on edge devices. Through extensive experiments and in-depth analysis, our benchmark reveals multiple important characteristics of token reduction in VPR and provides several practical insights into the trade-offs between accuracy and inference efficiency. For example, token reduction can reduce computational cost by up to 29\% and improve throughput by up to 44\%, while incurring less than 1\% degradation in recognition accuracy. Overall, this work establishes a comprehensive foundation for future research on token-efficient VPR and efficient visual retrieval systems. Our codes and models will be available at https://github.com/Tong-Jin01/TokenReduction4VPR
Problem

Research questions and friction points this paper is trying to address.

Visual Place Recognition
Token Reduction
Vision Transformers
Efficient Inference
Computational Overhead
Innovation

Methods, ideas, or system contributions that make the work stand out.

token reduction
visual place recognition
vision transformers
efficient inference
benchmark
T
Tong Jin
Shenyang Institute of Automation, Chinese Academy of Sciences; University of Chinese Academy of Sciences
Yunpeng Liu
Yunpeng Liu
Wuhan University of Technology
cement and concrete materials
S
Shuyu Hu
Shenyang Institute of Automation, Chinese Academy of Sciences; University of Chinese Academy of Sciences
Q
Qinghua Zhang
Shenyang Institute of Automation, Chinese Academy of Sciences; University of Chinese Academy of Sciences
Ruize Han
Ruize Han
SUAT
Computer VisionMultimedia AnalysisVideo UnderstandingActive Vision
S
Song Wang
Shenzhen University of Advanced Technology
F
Feng Lu
Shenzhen University of Advanced Technology