🤖 AI Summary
Image geolocation requires fine-grained reasoning by integrating multi-source contextual cues—visual, geographic, and cultural—yet existing methods suffer from low-quality training data and poor interpretability. To address this, we propose a language-driven analytical paradigm: (1) we construct NaviClues, the first high-quality, natural-language-annotated dataset derived from GeoGuessr; (2) we design a language-guided geographical reasoning mechanism that fuses multi-scale visual features with few-shot fine-tuning (<1,000 samples); and (3) we introduce a vision-language model (VLM) framework achieving a 14% reduction in mean localization error over standard benchmarks, significantly outperforming state-of-the-art methods. Our core innovation lies in explicitly leveraging natural language as an interpretable, structured reasoning medium—enabling high-accuracy, resource-efficient, and inherently explainable geolocation. The code and dataset are publicly released.
📝 Abstract
Image geo-localization is the task of predicting the specific location of an image and requires complex reasoning across visual, geographical, and cultural contexts. While prior Vision Language Models (VLMs) have the best accuracy at this task, there is a dearth of high-quality datasets and models for analytical reasoning. We first create NaviClues, a high-quality dataset derived from GeoGuessr, a popular geography game, to supply examples of expert reasoning from language. Using this dataset, we present Navig, a comprehensive image geo-localization framework integrating global and fine-grained image information. By reasoning with language, Navig reduces the average distance error by 14% compared to previous state-of-the-art models while requiring fewer than 1000 training samples. Our dataset and code are available at https://github.com/SparrowZheyuan18/Navig/.