Decoding the Disaster: Multi-Task Geospatial Reasoning with Vision-Language Models and Crowdsourced Imagery for Disaster Mapping

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of missing geospatial metadata, spatial ambiguity, and limited scalability in localizing crowdsourced disaster imagery by proposing the GRDisaster framework. Leveraging vision-language models (VLMs), this framework enables cross-view geo-localization and multi-view fusion. It presents the first systematic evaluation of VLMs’ spatial reasoning capabilities for disaster mapping, introducing a unified evaluation framework, interpretability verification metrics, and the PhotoMappers benchmark dataset. By effectively correlating crowdsourced images with reference imagery through structural and environmental cues, the approach accurately assesses disaster severity. Ultimately, this work transforms unstructured imagery into actionable geographic intelligence, significantly enhancing localization accuracy and output interpretability.
📝 Abstract
Crowdsourced imagery provides timely, fine-grained, street-level observations for disaster mapping, complementing conventional remote sensing imagery (RSI) during emergency response. However, such imagery is often unstructured, spatially ambiguous, and lacks reliable geographic metadata, making manual geolocalization and interpretation labor-intensive and difficult to scale. This work proposes a multi-task Geospatial Reasoning Disaster mapping framework, namely GRDisaster, to examine the potential of vision-language models (VLMs) in understanding, geolocalizing, and reasoning over crowdsourced disaster imagery. GRDisaster is built on a newly curated benchmark dataset derived from PhotoMappers, comprising 26,340 images organized into human-validated volunteered geographic information (VGI), street-view imagery (SVI), RSI cross-view triplets covering multiple disaster events from 2018 to 2024. The framework combines deterministic and probabilistic cross-view geolocalization with multi-view fusion to associate VGI images with georeferenced SVI and RSI. It introduces two sets of spatial reasoning indicators for cross-view geolocalization validation and disaster damage assessment. These indicators use structural, environmental, and global-scene cues to validate cross-view correspondences and visually observable damage evidence with expert-verified annotations to assess disaster severity, improving the interpretability of VLM outputs. To our knowledge, this study provides the first systematic investigation and unified evaluation framework for examining how VLM-based spatial reasoning can transform crowdsourced disaster imagery into actionable geospatial artificial intelligence (GeoAI) through cross-view geolocalization validation, interpretable spatial reasoning, and damage-aware severity assessment.
Problem

Research questions and friction points this paper is trying to address.

disaster mapping
crowdsourced imagery
geolocalization
spatial reasoning
GeoAI
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Models
Cross-view Geolocalization
Multi-task Geospatial Reasoning
Crowdsourced Imagery
Disaster Mapping
🔎 Similar Papers
No similar papers found.