[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119921-en":3,"doc-seo-119921-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119921,4810365810221,"Aurora","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Machine Learning for Gap-filling in Greenhouse Gas Emissions Databases - automated completion methods","Greenhouse Gas (GHG) emissions datasets often remain incomplete due to inconsistent reporting and limited transparency, which slows accurate assessment and hampers policy design for faster emissions reductions. This study assesses machine learning approaches to automate GHG database completion across three datasets of increasing complexity, comparing 18 gap-filling methods. Results indicate that simple interpolation is usually best for few available features or missing time steps, while richer feature sets and non-reporting emitters benefit from machine learning. Feature-importance outputs and scalable graph methods support prioritised data collection and ongoing updates.","1 Machine learning for gap-filling in greenhouse gas  \n2 emissions databases  \n3 Luke Cullen 1,* , Andrea Marinoni1,2 , and Jonathan Cullen 1  \n4 1 Department of Engineering, University of Cambridge, UK  \n5 2 Department of Physics and Technology, UiT the Arctic University of Norway, Norway  \n6 *[Corresponding author-lshc3@cam.ac.uk](Corresponding author-lshc3@cam.ac.uk)  \n7 ABSTRACT  \n8 Greenhouse Gas (GHG) emissions datasets are often incomplete due to inconsistent reporting and poor  \n9 transparency. Filling the gaps in these datasets allows for more accurate targeting of strategies aiming to  \n10 accelerate the reduction of GHG emissions. This study evaluates the potential of machine learning methods  \n11 to automate the completion of GHG datasets. We use 3 datasets of increasing complexity with 18 different  \n12 gap-filling methods and provide a guide to which methods are useful in which circumstances. If few dataset  \n13 features are available, or the gap consists only of a missing time step in a record, then simple interpolation  \n14 is often the most accurate method and complex models should be avoided. However, if more features are  \n15 available and the gap involves non-reporting emitters, then machine learning methods can be more accu- 16 rate than simple extrapolation. Furthermore, the secondary output of feature importance from complex  \n17 models allows for data collection prioritisation to accelerate the improvement of datasets. Graph based  \n18 methods are particularly scalable due to the ease of updating predictions given new data and incorporating  \n19 multimodal data sources. This study can serve as a guide to the community upon which to base ever more  \n20 integrated frameworks for automated detailed GHG emissions estimations, and implementation guidance  \n21 is available at [https://hackmd.io/@luke-scot/ML-for-GHG-database-completion](https://hackmd.io/@luke-scot/ML-for-GHG-database-completion and)[ and](https://hackmd.io/@luke-scot/ML-for-GHG-database-completion and)  \n22 [https://doi.org/10.5281/zenodo.10463104](https://doi.org/10.5281/zenodo.10463104) .  \n23  \n24 Keywords: Industrial Ecology, Greenhouse Gas Emissions, Automation, Machine Learning, Graph Represen-  \n25 tation Learning  \n26 1 INTRODUCTION  \n27 Greenhouse Gas (GHG) emissions datasets are often incomplete both at an individual facility level, e.g.  \n28 ClimateTRACE (2023), and at a national level, e.g. UNFCCC (2023) . Most countries, and many companies, 29 are accelerating their emissions reduction strategies inline with the Paris climate agreement and net-zero  \n30 objectives (Rogelj et al., 2016; Erb et al., 2022; Christiansen et al., 2023; Arnold and Toledano, 2021), but  \n31 incomplete and inaccurate datasets remain a barrier to understanding and therefore to effective policy-making  \n32 (IPCC, 2021; EPA, 2022a; Marlowe and Clarke, 2022) . In regions including the EU, UK and US, large 33 companies are required to use generic emissions intensity factors to convert their facilities’ activity data to 34 GHG emissions. These emissions factors are either provided by the relevant government authority such as 35 DEFRA and the EPA (DEFRA, 2009; EPA, 2022b), or obtained from Life Cycle Assessment (LCA) databases 36 including EcoInvent (Ecoinvent, 2022) . The facility-level data is then aggregated to a national estimate 37 grouped in to source types and reported yearly by UNFCCC Annex I parties.  \n38  \n39 Dataset incompleteness can originate from either the facility-level calculation or the national level aggregation.  \n40 At a facility level the 3 main causes of dataset incompleteness are: companies excluded from reporting  \n41 regulations, where data is simply not calculated; lack of transparency, sometimes as a result of industrial  \n42 secrecy; and non-compliant companies (Marlowe and Clarke, 2022; de Souza Leao et al., 2020) . Furthermore, 43 generic emissions factors used in calculations are non-specific to the production methods and supply ","cbCaicRNqMCTH5Al","https://ap.wps.com/l/cbCaicRNqMCTH5Al","pdf",1500359,1,22,"English","en",105,"# Abstract\n# Introduction\n## Sources of dataset incompleteness\n## Current approaches and limitations\n## Motivation for gap-filling using existing data","[{\"question\":\"Why are greenhouse gas emissions databases often incomplete?\",\"answer\":\"Incomplete coverage arises from multiple causes, including excluded reporting, lack of transparency, non-compliant reporting, and uncertainties introduced by generic emissions factors, plus inconsistencies in national reporting.\"},{\"question\":\"When does the study recommend simple interpolation over complex machine learning models?\",\"answer\":\"Interpolation is often most accurate when few dataset features are available or when the gap is only a missing time step in a record.\"},{\"question\":\"How can machine learning help when gaps involve non-reporting emitters?\",\"answer\":\"With more available features and gaps tied to non-reporting emitters, machine learning methods can outperform simple extrapolation and can provide feature-importance signals.\"}]","Machine Learning for Gap-filling in Greenhouse Gas Emissions Databases - automated completion methods | PDF",1785727004,55,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"machine-learning-for-gap-filling-in-greenhouse-gas-emissions-databases-automated-completion-methods","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/machine-learning-for-gap-filling-in-greenhouse-gas-emissions-databases-automated-completion-methods/119921/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why are greenhouse gas emissions databases often incomplete?","Question",{"text":75,"@type":76},"Incomplete coverage arises from multiple causes, including excluded reporting, lack of transparency, non-compliant reporting, and uncertainties introduced by generic emissions factors, plus inconsistencies in national reporting.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"When does the study recommend simple interpolation over complex machine learning models?",{"text":80,"@type":76},"Interpolation is often most accurate when few dataset features are available or when the gap is only a missing time step in a record.",{"name":82,"@type":73,"acceptedAnswer":83},"How can machine learning help when gaps involve non-reporting emitters?",{"text":84,"@type":76},"With more available features and gaps tied to non-reporting emitters, machine learning methods can outperform simple extrapolation and can provide feature-importance signals.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]