GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks

📅 2024-11-28
🏛️ arXiv.org
📈 Citations: 1
Influential: 0
📄 PDF

career value

176K/year
🤖 AI Summary
To address the suboptimal performance of general-purpose vision-language models (VLMs) on geospatial tasks—such as environmental monitoring and disaster response—this work introduces GeoBench, the first VLM benchmark dedicated to remote sensing understanding. GeoBench comprises over 10,000 human-verified, multi-source remote sensing instructions spanning six core tasks: scene understanding, fine-grained classification, object counting, localization, segmentation, and cross-temporal analysis. It systematically evaluates VLM capabilities on geospatially unique challenges, including detection of tiny objects, large-scale counting, and change identification. Experimental results reveal a significant capability gap: the state-of-the-art model LLaVA-OneVision achieves only 41.7% accuracy on multiple-choice tasks—substantially below its performance in general-domain benchmarks. GeoBench is publicly released to serve as a standardized evaluation platform for advancing geospatial AI research.

Technology Category

Application Category

📝 Abstract
While numerous recent benchmarks focus on evaluating generic Vision-Language Models (VLMs), they do not effectively address the specific challenges of geospatial applications. Generic VLM benchmarks are not designed to handle the complexities of geospatial data, an essential component for applications such as environmental monitoring, urban planning, and disaster management. Key challenges in the geospatial domain include temporal change detection, large-scale object counting, tiny object detection, and understanding relationships between entities in remote sensing imagery. To bridge this gap, we present GEOBench-VLM, a comprehensive benchmark specifically designed to evaluate VLMs on geospatial tasks, including scene understanding, object counting, localization, fine-grained categorization, segmentation, and temporal analysis. Our benchmark features over 10,000 manually verified instructions and spanning diverse visual conditions, object types, and scales. We evaluate several state-of-the-art VLMs to assess performance on geospatial-specific challenges. The results indicate that although existing VLMs demonstrate potential, they face challenges when dealing with geospatial-specific tasks, highlighting the room for further improvements. Notably, the best-performing LLaVa-OneVision achieves only 41.7% accuracy on MCQs, slightly more than GPT-4o, which is approximately double the random guess performance. Our benchmark is publicly available at https://github.com/The-AI-Alliance/GEO-Bench-VLM .
Problem

Research questions and friction points this paper is trying to address.

Evaluates VLMs for geospatial-specific challenges like temporal change detection.
Assesses VLMs on tasks including object counting, localization, and segmentation.
Highlights performance gaps in VLMs for geospatial applications.
Innovation

Methods, ideas, or system contributions that make the work stand out.

GEOBench-VLM for geospatial VLM evaluation
10,000+ verified geospatial instructions
Assesses VLMs on geospatial-specific tasks
🔎 Similar Papers
No similar papers found.
M
M. S. Danish
Mohamed bin Zayed University of Artificial Intelligence
Muhammad Akhtar Munir
Muhammad Akhtar Munir
Mohamed bin Zayed University of Artificial Intelligence, UAE
Deep LearningModel CalibrationDomain GeneralizationVLMsRemote Sensing
S
Syed Roshaan Ali Shah
University College London
K
Kartik Kuckreja
Mohamed bin Zayed University of Artificial Intelligence
F
F. Khan
Mohamed bin Zayed University of Artificial Intelligence, Linköping University, Sweden
P
Paolo Fraccaro
IBM Research Europe, UK
Alexandre Lacoste
Alexandre Lacoste
Staff Research Scientist, ServiceNow Research
machine learning
S
Salman Khan
Mohamed bin Zayed University of Artificial Intelligence, Australian National University