🤖 AI Summary
This study addresses the deficiency of existing vision-language models (VLMs) in kilometer-scale spatial intelligence and geographic layout understanding. Drawing upon cognitive science theories, this work introduces KilometerVision, the first benchmark constructed from real-world videos, alongside a multi-level spatial reasoning evaluation framework spanning from landmark anchoring to global map comprehension. The findings reveal a fundamental limitation of current VLMs: they circumvent complex path integration by relying on two-dimensional recognition and textual matching, thereby demonstrating that these models lack genuine spatial reasoning capabilities. To advance research in large-scale spatial intelligence, the proposed benchmark is made publicly available.
📝 Abstract
We push the frontier of large-scale spatial intelligence in Vision-Language Models (VLMs) and introduce the first benchmark that probes geographical layout understanding from real-world videos, spanning up to 1km distances. Inspired by the cognitive science literature, we evaluate models against the hierarchical stages of human spatial awareness: anchoring via landmarks, connecting them through routes, and integrating these into global mental maps. Extensive experiments reveal a fundamental divergence in how current AI models process spatial information. Instead of utilising true path integration or forming geometric survey knowledge, we find that VLMs rely almost entirely on 2D visual recognition and text-matching to bypass complex spatial reasoning. The benchmark is publicly available at https://perception-test-challenge.github.io/kilometervision.html.