π€ AI Summary
To address the latency of traditional data sources and the difficulty of extracting actionable information from social media during Canadian wildfire emergencies, this paper introduces WildFireCan-MMDβthe first multimodal social media dataset specifically designed for this context. It comprises real user-generated content (UGC) from X (formerly Twitter), meticulously annotated across 13 emergency response themes. Notably, it features the first Canada-localized, human-in-the-loop multimodal annotation effort for wildfire events. We propose a lightweight supervised paradigm (ViT + MLP) that achieves up to a 23% improvement in F1-score over zero-shot vision-language models (VLMs) for emergency theme classification. The dataset and code are publicly released to establish a reproducible benchmark and regionally adaptable methodology for emergency AI. This work empirically validates the critical importance of localized annotation and efficient fine-tuning in enhancing situational awareness during disasters.
π Abstract
Rapid information access is vital during wildfires, yet traditional data sources are slow and costly. Social media offers real-time updates, but extracting relevant insights remains a challenge. We present WildFireCan-MMD, a new multimodal dataset of X posts from recent Canadian wildfires, annotated across 13 key themes. Evaluating both Vision Language Models and custom-trained classifiers, we show that while zero-shot prompting offers quick deployment, even simple trained models outperform them when labelled data is available, by up to 23%. Our findings highlight the enduring importance of tailored datasets and task-specific training. Importantly, such datasets should be localized, as disaster response requirements vary across regions and contexts.