4MT-VLM: How Coarse Is a VLMs Cognitive Map?

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inability of vision-language models (VLMs) to maintain stable spatial cognitive maps under viewpoint changes. Drawing on clinical paradigms for probing hippocampal function, we construct a procedural landscape dataset that isolates pure spatial reasoning by disentangling appearance cues. Using a multimodal benchmark with a four-alternative forced-choice paradigm, we quantitatively evaluate cross-viewpoint recognition performance and spatial resolution granularity. Experiments reveal that state-of-the-art VLMs degrade to below chance-level performance following scene rotation. To our knowledge, this work is the first to systematically demonstrate that current VLMs lack robust three-dimensional spatial understanding, providing critical diagnostic insights for advancing future multimodal architectures.
📝 Abstract
An agent that moves must recognise a place from a viewpoint it has never seen. We introduce 4MT-VLM, a dataset of procedurally generated landscapes, each rendered across five stimulus modes that remove appearance cues while holding layout fixed: shape and colour, shape only, colour only, bare terrain peaks with no objects, and a valley viewpoint that puts the peaks on the horizon. The last condition is commonly used in clinics to probe hippocampal function in human patients. We test this benchmark across sixteen different open and closed-source models and report 4AFC performance, a measure which is also used to grade human participants. We observe that models identify a place from the studied viewpoint but lose it once the camera moves, dropping below the 25% chance level at 135° where a human observer scores 85%. Frontier models (Gemini 3.8 Flash, GPT-5.6) answer only 39% and 31% of rotated trials correctly, recovering to 85% and 55% only when distractors are moved more than 30 meters apart. Our benchmark demonstrates that while current VLMs possess rudimentary cognitive maps, their spatial resolution remains fundamentally too coarse to maintain a stable, 3D understanding of the world once the viewpoint changes.
Problem

Research questions and friction points this paper is trying to address.

Visual Language Models
Cognitive Map
Place Recognition
Viewpoint Change
Spatial Understanding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cognitive Map
4MT-VLM Benchmark
Visual Language Models
Spatial Reasoning
Procedural Generation