[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120267-en":3,"doc-seo-120267-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120267,13056703019404,"Miles","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Assessing models for de-identification of Electronic Discharge Summary using Machine Learning tools - Master of Science Project","De-identification removes identifying information from Electronic Discharge Summary clinical records to safeguard individual privacy and reduce misuse of personal data across collection, processing, distribution, and publication. The Electronic Discharge Summary contains Protected Health Information (PHI), making de-identification mandatory. This project applies machine learning to determine which model best de-identifies a dataset. Using an open-source Harvard Medical School dataset, Conditional Random Fields, LSTM, and Random Forest are evaluated on token-level F-measure, Recall, and Precision. LSTM achieves the highest micro averages and strong macro performance.","Assessing models for de-identification of Electronic Discharge Summary using Machine Learning tools  \nTshilisanani Mudau (11640828)  \nSupervisor(s):  \nProfessor Winston Garira  \nDr Rendani Netshikweta  \nA Project report submitted in partial fulfillment of the requirements for the degree of Master of Science in the field of e-Science  \nin the  \nDepartment of Computer Science and Applied Mathematics  \nUniversity of Venda  \n18 July 2024  \ni  \nDedicated to my late mom.  \nHow i wish you were here to witness God’s grace upon your only daughter.  \nii  \nDeclaration  \nI, Tshilisanani Mudau (11640828), declare that this report is my own, unaided work. It is being submitted for the degree of Master of Science in the field of e-Science atthe University of Venda. It has not been submitted for any degree or examination at any other university.  \nTshilisanani Mudau (11640828)  \n18 July 2024  \niii  \nAbstract  \nBackground: De-identification is a technique that eliminates identifying information from Clinical Records in order to protect individual privacy. This procedure decreases the chance of personal information being collected, processed, distributed, and published from being used to identify the person. When Machine Learning techniques were included in the de-identification process, it substantially improved over the previous method.  \nResearch Problem: The Electronic Discharge Summary(EDS) has evolved into a significantly improved technique of providing discharge summaries though this information contains Protected Health Information (PHI), which poses a risk to patients’ privacy. This makes the process of de-identification to be mandatory. There have lately been several Machine Learning approaches to de-identify data. This study focuses on applying Machine Learning techniques to figure out which model can best de-identify a data set.  \nMethods: The open source data set from Harvard Medical School was used. This data set contains 899 Electronic Health Records (EHR), 669 for training and 220 for test purpose. The Conditional Random Fields (CRF), Long Short Term Memory (LSTM) and Random Forest models were used, and the performance of each model was assessed.  \nFindings: In order to assess each model’s performance, evaluation metrics were used to compare F-measure, Recall and Precision at token level to determine which Machine Learning model performed best. The Long Short Term Memory was found to outperform both Conditional Random Fields and Random Forest with micro average F-measure, Recall and precision of 99%, and macro average F-measure of 77%, Recall of 73% and Precision of 90% .  \niv  \nAcknowledgements  \nAbove all, I would like to thank the Almighty God for everything, his unconditional love, mercy and grace. If it wasn’t of him, this would not have been possible.  \nI’d like to express my sincere gratitude to my supervisor, professor Winston Garira and my co-supervisor Dr Rendani Netshikweta for their assistance.  \nI would also like to thank everyone who assisted and supported me in any way.  \nMy heartfelt gratitude goes out to my family for their unfailing support.  \nFunding was provided by the DSI-NICIS NEPTTP and is highly appreciated.  \nv  \nContents  \nDeclaration ii  \nAbstract iii  \nAcknowledgements iv  \nList of Figures vii  \nList of Tables viii  \nList of Abbreviations ix  \n1 Introduction 1  \n1.1 Background ................................. 3  \n1.1.1 The purpose of De-identification ................. 3  \n1.1.2 Natural Language Processing (NLP) ............... 4  \n1.1.3 Privacy Requirement ........................ 5  \n1.1.4 Ethical Considerations ....................... 5  \n1.2 Problem Statement ............................. 5  \n1.3 Research Question ............................. 6  \n1.4 Research Aims and Objectives ....................... 6  \n1.4.1 Research Aims ........................... 6  \n1.4.2 Objectives .............................. 7  \n1.5 Limitations .................................. 7  \n1.6 Overview ....................","cbCaihGhMtq37Xtz","https://ap.wps.com/l/cbCaihGhMtq37Xtz","pdf",966127,1,58,"English","en",105,"# Introduction\n## Background\n## Problem Statement\n## Research Question\n## Research Aims and Objectives\n## Limitations\n## Overview\n# Literature Review\n## Related Work\n## Conclusion\n# Protection of Health Information in South Africa\n## How Personal Health Information is collected\n## How Personal Health Information is stored and protected\n## Logistics behind getting South African dataset\n## De-identification and its Rationale in USA\n## Sufficient degree of identification risk for a professional assessment\n## Validity period determined by experts for a given dataset\n# Research Methodology\n## Research design\n## Models\n## Data\n## Methods\n## Performance metrics\n## Analysis\n# Results and Discussion\n## Descriptive statistics\n## Experimental Results\n## Practical implications as a result of limitation\n# Conclusions and Future Work\n## Conclusions\n## Future Work","[{\"question\":\"Why is de-identification necessary for Electronic Discharge Summary records?\",\"answer\":\"Electronic Discharge Summary information includes Protected Health Information (PHI). De-identification is required to protect patient privacy and reduce the risk of personal information being used to identify individuals.\"},{\"question\":\"Which machine learning models were compared in the study?\",\"answer\":\"The project compares Conditional Random Fields (CRF), Random Forest, and Long Short-Term Memory (LSTM). Each model is assessed using token-level evaluation metrics.\"},{\"question\":\"How were the models evaluated, and which performed best?\",\"answer\":\"Models are evaluated using token-level F-measure, Recall, and Precision. LSTM outperformed CRF and Random Forest, reaching 99% micro-average scores and 77% macro-average F-measure, with Recall at 73% and Precision at 90%.\"}]","Assessing models for de-identification of Electronic Discharge Summary using Machine Learning tools - Master of Science Project | PDF",1785729144,146,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"assessing-models-for-de-identification-of-electronic-discharge-summary-using-machine-learning-tools-master-of-science-project","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/assessing-models-for-de-identification-of-electronic-discharge-summary-using-machine-learning-tools-master-of-science-project/120267/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is de-identification necessary for Electronic Discharge Summary records?","Question",{"text":75,"@type":76},"Electronic Discharge Summary information includes Protected Health Information (PHI). De-identification is required to protect patient privacy and reduce the risk of personal information being used to identify individuals.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which machine learning models were compared in the study?",{"text":80,"@type":76},"The project compares Conditional Random Fields (CRF), Random Forest, and Long Short-Term Memory (LSTM). Each model is assessed using token-level evaluation metrics.",{"name":82,"@type":73,"acceptedAnswer":83},"How were the models evaluated, and which performed best?",{"text":84,"@type":76},"Models are evaluated using token-level F-measure, Recall, and Precision. LSTM outperformed CRF and Random Forest, reaching 99% micro-average scores and 77% macro-average F-measure, with Recall at 73% and Precision at 90%.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]