[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-116953-en":3,"doc-seo-116953-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":20,"language":21,"language_code":22,"site_id":23,"html_lang":22,"table_of_contents":24,"faqs":25,"seo_title":26,"seo_description":14,"update_tm":27,"read_time":28},116953,549768064778,"Finn","https://ap-avatar.wpscdn.com/davatar_6f874abed73319feea01a86fa6f0fab8",8,"Research & Report","Machine Learning for Data Linkage","Data linkage compares individuals across sources using deterministic, probabilistic, or machine-learning approaches. This study evaluates traditional methods against selected machine learning models on a standard linkage task, comparing precision, recall, and F1 scores. Gradient boosted trees and a multiple layered perceptron classifier are supervised; maximum entropy classification is unsupervised. Models are trained and tested using a Census-to-CCS gold-standard linked dataset, with Splink (Fellegi-Sunter with Expectation Maximisation) as baseline.","International Journal of Population Data Science (2023) 8:3:059  \n\n| International Journal of\u003Cbr>Population Data Science\u003Cbr>Journal Website: [www.ijpds.org](www.ijpds.org) |  |\n| --- | --- |\n| Machine Learning for Data Linkage\u003Cbr>Rhosanna Ellum1 , Alex Lewis1 , Kristina Xhaferaj1 , Rachel Shipsey2 , Viktor Račinskij2 , and Zoe White2\u003Cbr>1 ONS, London, United Kingdom\u003Cbr>2 ONS, Titchfield, United Kingdom |  |\n\nData linkage traditionally uses deterministic and probabilistic methods. Alternatively, machine learning methods can be applied as classification algorithms, using the data to inform decisions. This project compared the quality, in terms of precision and recall, of traditional methods with selected machine learning methods when applied to a standard linkage problem.  \nTwo supervised methods, gradient boosted trees (GBT) and multiple layered perceptron classifier (MLPC), and one unsupervised method, maximum entropy classification (MEC), were implemented. The England and Wales 2021 Census to Census Coverage Survey (CCS) linkage was used as a goldstandard (GS) linked dataset to provide training samples for the supervised methods as well as testing samples for all methods. The F1 score (harmonic mean of precision and recall) was used to compare the performance of the models and to determine the optimal parameters and thresholds.  \nThe Splink implementation of Fellegi-Sunter with Expectation Maximisation was used as a baseline for comparison.  \nThe methods, trained on a sample of the GS, were used to link census and CCS data. All methods performed well with MEC achieving the highest precision (99 .79%) but lowest recall (96 .36%) . The MLPC model achieved the highest F1 score (98 .94%) .  \nTo understand the implications of not retraining supervised models for each dataset, the models were also used to link Census to a health dataset. The supervised models were not retrained using the health data; instead, the optimised GS models were applied. MEC had the lowest precision (96 .51%) but the highest recall (98 .48%) and highest F1 score (97 .49%) . With F1 scores of 96.99% and 96.14% respectively, the GBT and MLPC supervised models were not far behind in performance, despite not being trained using health data.  \nWe have shown that machine learning methods can be used effectively for data linkage problems. Unsurprisingly, supervised models perform best when trained on and applied to the same data. Further research into generic training may allow us to use both supervised and unsupervised machine learning models for future data linkage.  \n[https://doi.org/10.23889/ijpds.v8i2.2240](https://doi.org/10.23889/ijpds.v8i2.2240)  \nNovember 2023 © The Authors. Open Access under CC BY 4.0 ([https://creativecommons.org/licenses/by/4.0/deed.en](https://creativecommons.org/licenses/by/4.0/deed.en))","cbCaie0ozrDaNWLI","https://ap.wps.com/l/cbCaie0ozrDaNWLI","pdf",196807,1,"English","en",105,"# Background\n# Methods\n## Supervised models\n## Unsupervised model\n## Baseline comparison\n# Experiments and evaluation\n## Census to Census Coverage Survey\n## Census to health dataset without retraining\n# Results and conclusions","[{\"question\":\"Which machine learning methods were implemented for data linkage?\",\"answer\":\"The study implemented gradient boosted trees (GBT) and a multiple layered perceptron classifier (MLPC) as supervised methods, and maximum entropy classification (MEC) as an unsupervised method.\"},{\"question\":\"How was model performance compared across methods?\",\"answer\":\"Performance was compared using precision, recall, and the F1 score (the harmonic mean of precision and recall) to select optimal parameters and thresholds.\"},{\"question\":\"Did supervised models need retraining when linking to a different dataset?\",\"answer\":\"No. The supervised models were not retrained on the health dataset; instead, the optimized gold-standard models were applied directly.\"}]","Machine Learning for Data Linkage | PDF",1785672817,3,{"code":4,"msg":30,"data":31},"ok",{"site_id":23,"language":22,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"machine-learning-for-data-linkage","",{"@graph":35,"@context":84},[36,52,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,49],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":28},"https://docshare.wps.com/document/research-report/",{"item":50,"name":13,"@type":42,"position":51},"https://docshare.wps.com/document/machine-learning-for-data-linkage/116953/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":22,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-09-04","2026-08-02",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"Which machine learning methods were implemented for data linkage?","Question",{"text":74,"@type":75},"The study implemented gradient boosted trees (GBT) and a multiple layered perceptron classifier (MLPC) as supervised methods, and maximum entropy classification (MEC) as an unsupervised method.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How was model performance compared across methods?",{"text":79,"@type":75},"Performance was compared using precision, recall, and the F1 score (the harmonic mean of precision and recall) to select optimal parameters and thresholds.",{"name":81,"@type":72,"acceptedAnswer":82},"Did supervised models need retraining when linking to a different dataset?",{"text":83,"@type":75},"No. The supervised models were not retrained on the health dataset; instead, the optimized gold-standard models were applied directly.","https://schema.org",{"og:url":50,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":50},"index,follow",{"doc_id":7,"site_id":23},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":105,"slug":137},19,"General","general"]