[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120117-en":3,"doc-seo-120117-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120117,8796095461564,"Liam","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Machine learning for gap-filling in greenhouse gas emissions databases","Greenhouse gas (GHG) emissions datasets are often incomplete due to inconsistent reporting and limited transparency, which undermines evidence for strategies that accelerate emission reductions. This study assesses machine learning approaches to automate gap completion across three datasets with increasing complexity and tests 18 gap-filling methods. Simple interpolation is typically most accurate when features are scarce or gaps are confined to missing time steps. When additional features are available and gaps involve non-reporting emitters, machine learning can outperform extrapolation while providing feature-importance outputs that guide data collection priorities. Graph-based methods support scalable updating and integration of multimodal data sources, enabling more reliable, automated GHG estimation frameworks.","DOI: 10.1111/jiec.13507  \nMETHODS A RT ICLE  \nMachine learning for gap-filling in greenhouse gas emissions databases  \nLuke Cullen1   Andrea Marinoni1, 2  Jonathan Cullen1   \n1 Department of Engineering, University of Cambridge, Cambridge, UK  \n2 Department of Physics and Technology, UiT the Arctic University of Norway, Tromsø, Norway  \nCorrespondence  \nLuke Cullen, Department of Engineering, University of Cambridge, Cambridge, UK.  \nEmail: [lshc3@cam.ac.uk](lshc3@cam.ac.uk)  \nEditor Managing Review: Deepak Rajagopal  \nFunding information  \nUK Research and Innovation; UKRI Centre for Doctoral Training in Application of Artificial Intelligence to the study of Environmental Risks, Grant/Award Number: EP/S022961/1  \nAbstract  \nGreenhouse gas (GHG) emissions datasets are often incomplete due to inconsistent reporting and poor transparency. Filling the gaps in these datasets allows for more accurate targeting of strategies aiming to accelerate the reduction of GHG emissions. This study evaluates the potential of machine learning methods to automate the completion of GHG datasets. We use three datasets of increasing complexity with 18 different gap-filling methods and provide a guide to which methods are useful in which circumstances. If few dataset features are available, or the gap consists only of a missing time step in a record, then simple interpolation is often the most accurate method and complex models should be avoided. However, if more features are available and the gap involves non-reporting emitters, then machine learning methods can be more accurate than simple extrapolation. Furthermore, the secondary output of feature importance from complex models allows for data collection prioritization to accelerate the improvement of datasets. Graph-based methods are particularly scalable due to the ease of updating predictions given new data and incorporating multimodal data sources. This study can serve as a guide to the community upon which to base evermore integrated frameworks for automated detailed GHG emissions estimations, and implementation guidance is available at [https://hackmd.io/@luke-scot/ML-for-GHG](https://hackmd.io/@luke-scot/ML-for-GHG)database-completion and [https://doi.org/10.5281/zenodo.10463104. This article met](https://doi.org/10.5281/zenodo.10463104. This article met)[ ](https://doi.org/10.5281/zenodo.10463104. This article met)the requirements for a gold-gold JIE data openness badge described at [http://jie.click/](http://jie.click/)[ ](http://jie.click/)[badges.](badges.)  \nKEYWORDS  \nautomation, data completion, data prioritisation, graph representation learning, greenhouse gas emissions, machine learning  \nThis is an open access article under the terms of the Creative Commons Attribution License, which permits use, distribution and reproduction in any medium, provided the original work is properly cited.  \n© 2024 The Author(s). Journal of Industrial Ecology published by Wiley Periodicals LLC on behalf of International Society for Industrial Ecology.  \nCULLEN ET AL.  \n1  INTRODUCTION  \nGreenhouse gas (GHG) emissions datasets are often incomplete both at an individual facility level, for example, ClimateTRACE (2023), and at a national level, for example, UNFCCC (2023). Most countries, and many companies, are accelerating their emissions reduction strategies in line with the Paris climate agreement and net-zero objectives (Arnold & Toledano, 2021; Christiansen et al., 2023; Erb et al., 2022; Rogelj et al., 2016), but incomplete and inaccurate datasets remain a barrier to understanding and, therefore, to effective policy-making(EPA,2022a;IPCC,2021;Marlowe &Clarke, 2022). In regions including the European Union, the United Kingdom, and the United States, large companies are required to use generic emissions intensity factors to convert their facilities’ activity data to GHG emissions. These emissions factors are either provided by the relevant government authority such as DEFRA and the EPA (DEFRA, 2009;E","cbCailZ0zBLVko6U","https://ap.wps.com/l/cbCailZ0zBLVko6U","pdf",1407651,1,12,"English","en",105,"# Introduction\n## Origins of dataset incompleteness\n## Impacts on policy and emissions calculations\n# Machine learning approach for gap filling\n## Method comparison across datasets\n## When interpolation works best\n## When machine learning outperforms extrapolation\n# Feature importance and prioritizing data collection\n## Using secondary outputs from complex models\n# Scalability with graph-based methods\n## Updating predictions with new data\n## Integrating multimodal data sources","[{\"question\":\"Why are greenhouse gas emissions datasets often incomplete?\",\"answer\":\"They are frequently incomplete due to inconsistent reporting and poor transparency. Gaps can arise from facility-level calculations, national-level aggregation, missing computations, or non-reporting and non-compliant reporting.\"},{\"question\":\"Which gap-filling methods are most accurate when features are limited?\",\"answer\":\"When few dataset features are available or the gap is only a missing time step, simple interpolation is often the most accurate. Complex models should generally be avoided in these situations.\"},{\"question\":\"When can machine learning outperform simple extrapolation?\",\"answer\":\"Machine learning can be more accurate when more features are available and the gap involves non-reporting emitters rather than only missing time steps. The study also highlights the value of feature-importance outputs for guiding data collection.\"}]","Machine learning for gap-filling in greenhouse gas emissions databases | PDF",1785728292,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"machine-learning-for-gap-filling-in-greenhouse-gas-emissions-databases","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/machine-learning-for-gap-filling-in-greenhouse-gas-emissions-databases/120117/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why are greenhouse gas emissions datasets often incomplete?","Question",{"text":75,"@type":76},"They are frequently incomplete due to inconsistent reporting and poor transparency. Gaps can arise from facility-level calculations, national-level aggregation, missing computations, or non-reporting and non-compliant reporting.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which gap-filling methods are most accurate when features are limited?",{"text":80,"@type":76},"When few dataset features are available or the gap is only a missing time step, simple interpolation is often the most accurate. Complex models should generally be avoided in these situations.",{"name":82,"@type":73,"acceptedAnswer":83},"When can machine learning outperform simple extrapolation?",{"text":84,"@type":76},"Machine learning can be more accurate when more features are available and the gap involves non-reporting emitters rather than only missing time steps. The study also highlights the value of feature-importance outputs for guiding data collection.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]