NAVIG: Natural Language-guided Analysis with Vision Language Models for Image Geo-localization

📅 2025-02-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Image geolocation requires fine-grained reasoning by integrating multi-source contextual cues—visual, geographic, and cultural—yet existing methods suffer from low-quality training data and poor interpretability. To address this, we propose a language-driven analytical paradigm: (1) we construct NaviClues, the first high-quality, natural-language-annotated dataset derived from GeoGuessr; (2) we design a language-guided geographical reasoning mechanism that fuses multi-scale visual features with few-shot fine-tuning (<1,000 samples); and (3) we introduce a vision-language model (VLM) framework achieving a 14% reduction in mean localization error over standard benchmarks, significantly outperforming state-of-the-art methods. Our core innovation lies in explicitly leveraging natural language as an interpretable, structured reasoning medium—enabling high-accuracy, resource-efficient, and inherently explainable geolocation. The code and dataset are publicly released.

Technology Category

Computer Vision: Language and VisionNatural Language Processing: Language Grounding & Multi-modal NLPKnowledge Representation and Reasoning: Geometric, Spatial, and Temporal Reasoning

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
Image geo-localization is the task of predicting the specific location of an image and requires complex reasoning across visual, geographical, and cultural contexts. While prior Vision Language Models (VLMs) have the best accuracy at this task, there is a dearth of high-quality datasets and models for analytical reasoning. We first create NaviClues, a high-quality dataset derived from GeoGuessr, a popular geography game, to supply examples of expert reasoning from language. Using this dataset, we present Navig, a comprehensive image geo-localization framework integrating global and fine-grained image information. By reasoning with language, Navig reduces the average distance error by 14% compared to previous state-of-the-art models while requiring fewer than 1000 training samples. Our dataset and code are available at https://github.com/SparrowZheyuan18/Navig/.
Problem

Research questions and friction points this paper is trying to address.

Improving image geo-localization accuracy
Creating high-quality dataset for reasoning
Reducing distance error with fewer samples
Innovation

Methods, ideas, or system contributions that make the work stand out.

Integrates global and fine-grained image information
Uses language-guided reasoning for geo-localization
Reduces distance error with fewer training samples
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zheyuan Zhang
Tsinghua University
R
Runze Li
Nanjing University
Tasnim Kabir
Tasnim Kabir
University of Maryland-- College Park
Natural Language ProcessingMachine Learning
J
Jordan L. Boyd-Graber
University of Maryland