WildFireCan-MMD: A Multimodal dataset for Classification of User-generated Content During Wildfires in Canada

πŸ“… 2025-04-17
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
To address the latency of traditional data sources and the difficulty of extracting actionable information from social media during Canadian wildfire emergencies, this paper introduces WildFireCan-MMDβ€”the first multimodal social media dataset specifically designed for this context. It comprises real user-generated content (UGC) from X (formerly Twitter), meticulously annotated across 13 emergency response themes. Notably, it features the first Canada-localized, human-in-the-loop multimodal annotation effort for wildfire events. We propose a lightweight supervised paradigm (ViT + MLP) that achieves up to a 23% improvement in F1-score over zero-shot vision-language models (VLMs) for emergency theme classification. The dataset and code are publicly released to establish a reproducible benchmark and regionally adaptable methodology for emergency AI. This work empirically validates the critical importance of localized annotation and efficient fine-tuning in enhancing situational awareness during disasters.

Technology Category

Computer Vision: Multi-modal VisionMachine Learning: Large Multimodal Models (LMMs)Data Mining & Knowledge Management: Mining of Visual, Multimedia & Multimodal Data

Application Category

Web Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web dataSearch and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGSocial Networks and Social Media: Generative AI / large language models and their impact on social systems
πŸ“ Abstract
Rapid information access is vital during wildfires, yet traditional data sources are slow and costly. Social media offers real-time updates, but extracting relevant insights remains a challenge. We present WildFireCan-MMD, a new multimodal dataset of X posts from recent Canadian wildfires, annotated across 13 key themes. Evaluating both Vision Language Models and custom-trained classifiers, we show that while zero-shot prompting offers quick deployment, even simple trained models outperform them when labelled data is available, by up to 23%. Our findings highlight the enduring importance of tailored datasets and task-specific training. Importantly, such datasets should be localized, as disaster response requirements vary across regions and contexts.
Problem

Research questions and friction points this paper is trying to address.

Classifying wildfire-related social media content effectively
Comparing zero-shot models vs trained classifiers for accuracy
Addressing regional variability in disaster response data needs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal dataset for wildfire content classification
Vision Language Models vs custom-trained classifiers
Localized datasets enhance disaster response accuracy
πŸ”Ž Similar Papers
No similar papers found.