[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86234-en":3,"doc-seo-86234-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86234,1374391974585,"Genevieve","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","A Multimodal Dataset for Large Language Model Applications in the Energy Domain","mAIEnergy is an open-access, multimodal corpus designed to enable Large Language Model (LLM) applications in the energy sector. The dataset combines about 50,000 textual documents, 20,000 images, 25 million numerical time-series records, and 2 million geospatial and relational entries. It covers policy and regulatory materials, scientific literature, news, satellite and contextual imagery, electricity measurements, weather observations, statistics, and structured representations of energy infrastructure. Harmonized outputs include consistent metadata and reproducible retrieval workflows following FAIR principles, supporting AI-driven energy research, modeling, and decision-making.","arXiv :2607 . 11459v1 [ ee ss . SY] 13 Jul 2026  \nA Multimodal Dataset for Large Language Model Applications in the Energy Domain  \nCostas Mylonas∗1 and Magda Foti†1  \n1 Energy Digitalization Group, UBITECH, Athens, Greece  \nAbstract  \nThis paper presents the mAIEnergy dataset, an open-access, multimodal corpus developed to support Large Language Model (LLM) applications in the energy sector.  \nThe dataset integrates approximately 50,000 textual documents, 20,000 images, 25 million numerical time series records, and 2 million geospatial and relational data entries. It includes policy and regulatory texts, scientific articles and news articles, satellite and contextual imagery, electricity system measurements, weather observations, statistical indicators, and geospatial representations of energy infrastructure and related entities. All data have been harmonized into structured, ready-to-use formats, accompanied by consistent metadata and reproducible data retrieval and preparation workflows. The dataset can serve as a foundational energy knowledge base, allowing energy stakeholders to integrate additional open-source or proprietary data. The mAIEnergy dataset adheres to Findable, Accessible, Interoperable, and Reusable (FAIR) principles, enhancing its applicability for AI-driven energy research, modeling, and decision-making.  \nBackground & Summary  \nThe energy sector is undergoing rapid transformation driven by ambitious sustainability targets, decarbonization initiatives, and increasing digitalization. Effectively navigating these changes relies on advanced analytical tools capable of processing diverse and complex energy-related data. Artificial Intelligence (AI), particularly through Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG), has shown significant potential for energy-focused tasks, including information retrieval, knowledge integration, and decision-support applications [1] . However, the effectiveness of these AI solutions is highly dependent on access to AI-ready, multimodal, and domain-specific datasets.  \nCurrently, publicly available energy datasets typically focus on single modalities. For example, numerical data are provided from sources like the European Network of Transmission System Operators for Electricity (ENTSO-E) Transparency Platform [2 , 3] or Eurostat statistics [4] . Geospatial data come from infrastructure mapping services, such  \n∗[kmylonas@ubitech.eu](kmylonas@ubitech.eu)[ ](kmylonas@ubitech.eu)†[mfoti@ubitech.eu](mfoti@ubitech.eu)  \nas OpenStreetMap [5], while textual information is distributed across policy and research documents. Similarly, satellite imagery is sourced from platforms like the European Copernicus [6] . This fragmented landscape constrains their applicability for developing advanced AI-driven decision-support tools and conversational assistants, which rely on integrated multimodal datasets to achieve contextual understanding and reasoning capabilities [7] .  \nFurthermore, the adoption of Open and Findable, Accessible, Interoperable, Reusable (FAIR) data practices, as defined by the relevant guiding principles [8], is becoming increasingly essential in the research community, including energy research. Nevertheless, the energy sector has been found to lag behind other domains in adopting comprehensive open-data and open-source practices [9] . Bridging this gap requires the creation of harmonized, multimodal datasets that comply with FAIR principles and foster transparency, accessibility, and interoperability across the energy research ecosystem.  \nIn response to this need, we introduce the mAIEnergy dataset [10], an open-access corpus specifically designed to advance LLM-driven applications and research in the energy domain. The dataset integrates four distinct data modalities, namely textual, imagery, numerical, and geospatial to provide a comprehensive representation of the energy landscape.  \nTextual sources include articles from Wikipedia [11], news art","cbCaikDJpFr75gIP","https://ap.wps.com/l/cbCaikDJpFr75gIP","pdf",867822,3,1,27,"English","en",105,"# Abstract\n# Background & Summary\n# Dataset Overview\n## Textual Sources\n## Imagery Sources\n## Numerical Data\n## Geospatial and Relational Data","[{\"question\":\"What is the mAIEnergy dataset and what is it intended for?\",\"answer\":\"mAIEnergy is an open-access multimodal corpus created to support LLM-driven applications and research in the energy domain, enabling integrated knowledge and decision support.\"},{\"question\":\"What modalities and approximate scale does mAIEnergy include?\",\"answer\":\"It integrates around 50,000 textual documents, 20,000 images, 25 million numerical time-series records, and about 2 million geospatial and relational data entries.\"},{\"question\":\"How does the dataset help overcome limitations of existing public energy datasets?\",\"answer\":\"It addresses fragmentation by harmonizing multiple modalities—text, imagery, numerical, and geospatial—into structured, ready-to-use formats for contextual understanding and reasoning.\"}]",1784209689,68,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"a-multimodal-dataset-for-large-language-model-applications-in-the-energy-domain","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/a-multimodal-dataset-for-large-language-model-applications-in-the-energy-domain/86234/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the mAIEnergy dataset and what is it intended for?","Question",{"text":75,"@type":76},"mAIEnergy is an open-access multimodal corpus created to support LLM-driven applications and research in the energy domain, enabling integrated knowledge and decision support.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What modalities and approximate scale does mAIEnergy include?",{"text":80,"@type":76},"It integrates around 50,000 textual documents, 20,000 images, 25 million numerical time-series records, and about 2 million geospatial and relational data entries.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the dataset help overcome limitations of existing public energy datasets?",{"text":84,"@type":76},"It addresses fragmentation by harmonizing multiple modalities—text, imagery, numerical, and geospatial—into structured, ready-to-use formats for contextual understanding and reasoning.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]