[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-125253-en":3,"doc-seo-125253-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},125253,687197207057,"Sage","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Machine Learning with Administrative Data for Energy Poverty Identification in the UK","Energy poverty remains a critical challenge requiring efficient, scalable identification to enable targeted interventions. UK monitoring has relied on the Low Income Low Energy Efficiency (LILEE) and, previously, the Low Income High Costs (LIHC) indicators, but their reliance on complex, time-intensive data collection limits their use for pinpointing specific households. This study builds machine learning classifiers trained on administrative-data-like inputs using English Housing Survey variables capturing household socio-demographics and building characteristics, addressing class imbalance via resampling and class weighting, and comparing against a UK government benchmark. Performance is evaluated with accuracy, balanced accuracy, precision, recall, and F1, with SHAP values supporting interpretability; the best XGBoosting model improves balanced accuracy and precision, highlighting income and dwelling characteristics as key determinants.","Article  \nMachine Learning with Administrative Data for Energy Poverty Identification in the UK  \nLin Zheng  and Eoghan McKenna *  \nAcademic Editor: Jin-Li Hu  \nReceived: 24 March 2025  \nRevised: 3 June 2025  \nAccepted: 5 June 2025  \nPublished: 9 June 2025  \nCitation: Zheng, L.; McKenna, E. Machine Learning with Administrative Data for Energy Poverty Identification in the UK. Energies 2025, 18, 3054. [https://](https://)[ ](https://)[doi.org/10.3390/en18123054](doi.org/10.3390/en18123054)  \n[Copyright:](Copyright:) © 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license ([https://creativecommons.org/lice](https://creativecommons.org/lice)[nses/by/4.0/](nses/by/4.0/)) .  \nUCL Energy Institute, University College London, 14 Upper Woburn Place, London WC1H 0NN, UK; [lin.z@ucl.ac.uk](lin.z@ucl.ac.uk)  \n* Correspondence: [e.mckenna@ucl.ac.uk](e.mckenna@ucl.ac.uk)  \nAbstract: Energy poverty continues to be a critical challenge, and this requires efficient and scalable identification methods to support targeted interventions. The Low Income Low Energy Efficiency (LILEE) indicator and previously the Low Income High Costs (LIHC) indicator have been used by the UK government to monitor national energy poverty levels. Yet due to their reliance on complex, time-intensive data collection processes and estimations, these indicators are not suitable for identifying energy poverty in specific households. This study investigates an alternative approach to energy poverty identification: using machine learning models trained on administrative data, data that could reasonably be available to governments for all or most households. We develop machine learning models using data from the English Housing Survey that serves as a proxy for administrative data. This data is selected to closely resemble what might be available in national administrative databases, incorporating variables such as household socio-demographics and building physical characteristics. We evaluate multiple classification algorithms, including Random Forest and XGBoosting, applying resampling and class weighting techniques to address the inherent class imbalance in energy poverty classification. We compare model performance with a ‘benchmark’ model developed by the UK government for the same goal. Model performance is assessed using the metrics of accuracy, balanced accuracy, precision, recall, and F1-score, with SHapley Additive exPlanations (SHAP) values providing the interpretability of the predictions. The best-performing model (XGBoosting with class weighting) achieves higher balanced accuracy (0.88), and precision (0.51) compared to the benchmark model (balanced accuracy: 0.77, precision: 0.24), demonstrating an improved ability to classify energy-poor households with fewer data constraints. SHAP analysis reveals household income and dwelling characteristics are key determinants of energy poverty. This research demonstrates that machine learning, trained on existing administrative datasets, offers a feasible, scalable, and interpretable alternative for energy poverty identification, enabling new opportunities for efficient targeted policy interventions. This study also aligns with recent UK government discussions on the potential for integrating administrative data sources to enhance policy implementation. Future research could explore the integration of real-time smart meter data to refine energy poverty assessments further.  \nKeywords: energy poverty; machine learning; administrative data; SHAP value; classification algorithms  \n1. Introduction  \nEnergy poverty is a critical issue that affects households’ ability to afford adequate energy for heating, cooling, and other essential services [1–3], also called “fuel poverty”.  \nResidents that are energy-poor have high risk of health problems such as respiratory infections and worsened chron","cbCaihdJ6NHt2CgX","https://ap.wps.com/l/cbCaihdJ6NHt2CgX","pdf",1244450,1,26,"English","en",105,"# Introduction\n## The Scope of Energy Poverty in England\n## The Need for Alternative Approaches","[{\"question\":\"Why are the UK’s existing LILEE and LIHC indicators insufficient for identifying energy-poor households?\",\"answer\":\"They depend on complex, time-intensive data collection and multi-step estimations, which makes them costly and less suitable for targeting specific households.\"},{\"question\":\"What approach does the study propose for energy poverty identification?\",\"answer\":\"It develops machine learning models trained on administrative-data-like variables that could plausibly be available in national government records.\"},{\"question\":\"Which model performs best and how is interpretability handled?\",\"answer\":\"The best-performing model is XGBoosting with class weighting, evaluated using standard classification metrics, and interpretability is provided through SHAP values.\"}]","Machine Learning with Administrative Data for Energy Poverty Identification in the UK | PDF",1785897737,66,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"machine-learning-with-administrative-data-for-energy-poverty-identification-in-the-uk","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/machine-learning-with-administrative-data-for-energy-poverty-identification-in-the-uk/125253/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why are the UK’s existing LILEE and LIHC indicators insufficient for identifying energy-poor households?","Question",{"text":75,"@type":76},"They depend on complex, time-intensive data collection and multi-step estimations, which makes them costly and less suitable for targeting specific households.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What approach does the study propose for energy poverty identification?",{"text":80,"@type":76},"It develops machine learning models trained on administrative-data-like variables that could plausibly be available in national government records.",{"name":82,"@type":73,"acceptedAnswer":83},"Which model performs best and how is interpretability handled?",{"text":84,"@type":76},"The best-performing model is XGBoosting with class weighting, evaluated using standard classification metrics, and interpretability is provided through SHAP values.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]