🤖 AI Summary
This study addresses the limitations of unidirectional sampling, which often misses critical cues, and hierarchical prediction conflicts in complex multi-shot video geolocalization. To this end, we propose a method integrating bidirectional temporal sampling with a hierarchical Graph Attention Network (GAT). By constructing complementary spatiotemporal features and introducing a structure-aware message passing mechanism, our approach effectively eliminates hierarchical conflicts. Furthermore, a dual-constraint pruning strategy is designed to enhance inference efficiency. We also release GeoGAT10k, a benchmark dataset containing multi-shot videos, to validate generalization capabilities. Experimental results demonstrate that the proposed method achieves state-of-the-art performance on both CityGuessr68k and GeoGAT10k, completely resolving hierarchical conflicts while improving city-level accuracy by 2.6% and over 24%, respectively.
📝 Abstract
Global video geo-localization aims to infer the geographic location of a video worldwide, evaluating performance across four geographic hierarchies: city, state/province, country, and continent. Existing methods typically employ one-way uniform sampling to process video frames and train independent classifiers for each hierarchy, which leads to the loss of key geographic cues and prediction conflicts between hierarchies, especially for complex multi-shot edited videos. To address these limitations, we propose GeoGAT, which integrates bidirectional temporal sampling with graph attention networks (GATs). Specifically, GeoGAT extracts forward and offset-reversed frame sequences to construct complementary spatiotemporal features. These fused features are then fed into a predefined geographical hierarchy graph, where GATs perform structure-aware message passing, while a dual-constraint mechanism prunes predictions to eliminate cross-hierarchy conflicts. We construct GeoGAT10k, comprising 9,720 multi-shot edited videos from 166 cities worldwide, specifically to benchmark generalization ability on complex video structures. Experimental results on CityGuessr68k and GeoGAT10k demonstrate that GeoGAT eliminates hierarchical conflicts entirely and achieves state-of-the-art performance across all four geographic hierarchies. On CityGuessr68k, GeoGAT outperforms the strongest baseline, evaluated under both classification and retrieval protocols, by 2.6 percentage points at the city level. On the more challenging GeoGAT10k with multi-shot edited videos, the accuracy improvement exceeds 24 percentage points, validating strong generalization to complex real-world scenarios.