🤖 AI Summary
This study addresses the critical scarcity of high-quality, multimodal, and structured open datasets in the energy domain that are essential for advancing large language model applications. To bridge this gap, the authors introduce mAIEnergy, the first comprehensive multimodal energy dataset integrating heterogeneous sources—including policy documents, scientific literature, power system records, meteorological data, and infrastructure information—across four modalities: text, images, time series, and geospatial data. Through standardized preprocessing, unified metadata management, and strict adherence to FAIR principles, mAIEnergy provides a consistent, reproducible corpus comprising 50,000 textual documents, 20,000 images, 25 million time-series entries, and 2 million geospatial records. This resource establishes a foundational infrastructure for AI-driven energy research and decision-making.
📝 Abstract
This paper presents the mAIEnergy dataset, an open-access, multimodal corpus developed to support Large Language Model (LLM) applications in the energy sector. The dataset integrates approximately 50,000 textual documents, 20,000 images, 25 million numerical time series records, and 2 million geospatial and relational data entries. It includes policy and regulatory texts, scientific articles and news articles, satellite and contextual imagery, electricity system measurements, weather observations, statistical indicators, and geospatial representations of energy infrastructure and related entities. All data have been harmonized into structured, ready-to-use formats, accompanied by consistent metadata and reproducible data retrieval and preparation workflows. The dataset can serve as a foundational energy knowledge base, allowing energy stakeholders to integrate additional open-source or proprietary data. The mAIEnergy dataset adheres to Findable, Accessible, Interoperable, and Reusable (FAIR) principles, enhancing its applicability for AI-driven energy research, modeling, and decision-making.