MultiGlobeQA: A Multilingual and Globally Diverse Benchmark for Geospatial Reasoning

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of current large language models (LLMs) in performing geometric and topological computations for geospatial reasoning and the absence of a multilingual, globally representative evaluation benchmark. We introduce GeoBench, a comprehensive geospatial reasoning benchmark covering 201 countries and regions, supporting 17 languages, and comprising 46,060 questions that span 14 spatial functions and 15 answer formats, with executable ground truth derived from a knowledge graph. Through stratified sampling by income and population density, multilingual alignment, and diverse LLM evaluation paradigms, our analysis reveals fundamental bottlenecks in models’ computational capabilities—such as grid indexing and shape operations—with overall accuracy remaining below 66% even when provided with gold-standard facts. Notably, performance disparities worsen for low-income regions after fact injection, highlighting persistent inequities in model generalization.
📝 Abstract
Geospatial reasoning, i.e., computing distances, containment, and other spatial relations over real-world entities, is central to navigation and logistics, yet large language models (LLMs) struggle with the required geometric and topological computation despite storing considerable geographic knowledge. Existing benchmarks localize these failures only partially: they are synthetic or smallscale, largely monolingual, and offer limited control over geographic coverage. We introduce MultiGlobeQA, a multilingual benchmark of 46,060 question-answer pairs spanning 14 spatial-function families and 15 answer formats, with execution-based ground truth over three knowledge graphs. It covers 201 countries and territories via income- and density-stratified sampling, with parallel questions in English and 16 additional high- and low-resource languages. Across parametric, reasoning, and agentic settings, LLMs collapse on tasks requiring grid indexing and shape computation, while topological relations and directions fare best. Retrieval and tool use yield considerable gains, yet performance plateaus below two thirds even when gold facts are supplied, indicating that computation, not access to knowledge, is the bottleneck. Models also underperform on low-income regions, a gap that gold facts widen rather than close.
Problem

Research questions and friction points this paper is trying to address.

geospatial reasoning
large language models
multilingual benchmark
spatial relations
geographic coverage
Innovation

Methods, ideas, or system contributions that make the work stand out.

geospatial reasoning
multilingual benchmark
knowledge graph
stratified sampling
large language models
🔎 Similar Papers
No similar papers found.