[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123609-en":3,"doc-seo-123609-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123609,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1785132997149421697",8,"Research & Report","Sensitive Data Detection with High-Throughput Machine Learning Models in Electrical Health Records","Big data drives increasing demands for sharing healthcare information to improve health outcomes and advance research, yet HIPAA requires protecting protected health information (PHI). Traditional de-identification workflows lack efficient, portable tools for detecting PHI, especially because PHI fields vary heterogeneously across organizations and datasets. This work leverages machine learning on structured EHR data using engineered metadata features to distinguish PHI from non-PHI fields, achieving 99% accuracy on unseen datasets.","Sensitive Data Detection with High-Throughput Machine Learning Models in Electrical  \nHealth Records  \nKai Zhang, PhD, Xiaoqian Jiang, PhD  \nThe University of Texas Health Science Center,  \nMcWilliams School of Biomedical Informatics, Houston, TX, USA  \nAbstract:  \nIn the era of big data, there is an increasing need for healthcare providers, communities, and researchers to share data and collaborate to improve health outcomes, generate valuable insights, and advance research. The Health Insurance Portability and Accountability Act of 1996 (HIPAA) is a federal law designed to protect sensitive health information by defining regulations for protected health information (PHI) . However, it does not provide efficient tools for detecting or removing PHI before data sharing. One of the challenges in this area of research is the heterogeneous nature of PHI fields in data across different parties. This variability makes rule-based sensitive variable identification systems that work on one database fail on another. To address this issue, our paper explores the use of machine learning algorithms to identify sensitive variables in structured data, thus facilitating the de-identification process. We made a key observation that the distributions of metadata of PHI fields and non-PHI fields are very different. Based on this novel finding, we engineered over 30 features from the metadata of the original features and used machine learning to build classification models to automatically identify PHI fields in structured Electronic Health Record (EHR) data. We trained the model on a variety of large EHR databases from different data sources and found that our algorithm achieves 99% accuracy when detecting PHI-related fields for unseen datasets. The implications of our study are significant and can benefit industries that handle sensitive data.  \nKeywords: De-identification, Protected health information (PHI), Electronic health records (EHR), Machine learning algorithms  \n1. Introduction  \nBecause of improvements in online data tracking and sharing techniques, healthcare data privacy has become a major issue in recent years. In the fields of medicine and research, the sharing of data with other parties and for secondary uses is a frequent practice 1. Patients, however, have voiced worries regarding the lack of control they have over how their data is used and shared 2. There are many difficulties in maintaining data confidentiality in the context of health care research3. Although there are legal ways to share data, there is always space for improvement in terms of speeding up and securing data transmissions. To encourage the safe and responsible use of healthcare data, it is critical to address these concerns as soon as they arise.  \nDe-identification is the process of taking personal data out of a dataset so that it cannot be connected to particular people4. This procedure is crucial for safeguarding the privacy of people whose personal data is present in datasets. Personal information is described as data that can be used to identify an individual5. In the healthcare sector, where patient data sets frequently contain sensitive information, de-identification is especially crucial6. The Health Insurance Portability and Accountability Act's (HIPAA) Privacy Regulation offers advice on the best ways to achieve deidentification while still adhering to HIPAA rules7. De-identifying datasets requires considering both direct and indirect identifiers. In contrast to indirect identifiers, which may not be sufficient on their own to identify an individual but can result in identification when paired with other information, direct identifiers are sufficient alone to potentially identify an individual 8. For the protection of structured data, a variety of techniques and resources are available. They include cryptographic techniques that can safeguard patient data while upholding HIPAA compliance and preserving data links 9. Another method for protecting pri","cbCailybEalYDIPi","https://ap.wps.com/l/cbCailybEalYDIPi","pdf",721501,1,17,"English","en",105,"# Introduction\n## Privacy challenges in healthcare data sharing\n## De-identification concepts and HIPAA context\n## Common de-identification methods and tools","[{\"question\":\"Why are rule-based PHI detection systems difficult to generalize across datasets?\",\"answer\":\"PHI fields are heterogeneous across different parties, so patterns learned in one database may not hold in another. This variability can cause rule-based identification to fail on unseen datasets.\"},{\"question\":\"What metadata-driven strategy does the approach use to identify PHI fields?\",\"answer\":\"The method observes that the metadata distributions of PHI fields and non-PHI fields differ. It engineers features from original feature metadata and trains classification models to identify PHI-related variables.\"},{\"question\":\"How well does the proposed machine learning algorithm perform on unseen datasets?\",\"answer\":\"Training on multiple large EHR databases from different sources, the algorithm achieves 99% accuracy for detecting PHI-related fields on datasets not seen during training.\"}]","Sensitive Data Detection with High-Throughput Machine Learning Models in Electrical Health Records | PDF",1785817621,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"sensitive-data-detection-with-high-throughput-machine-learning-models-in-electrical-health-records","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/sensitive-data-detection-with-high-throughput-machine-learning-models-in-electrical-health-records/123609/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why are rule-based PHI detection systems difficult to generalize across datasets?","Question",{"text":75,"@type":76},"PHI fields are heterogeneous across different parties, so patterns learned in one database may not hold in another. This variability can cause rule-based identification to fail on unseen datasets.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What metadata-driven strategy does the approach use to identify PHI fields?",{"text":80,"@type":76},"The method observes that the metadata distributions of PHI fields and non-PHI fields differ. It engineers features from original feature metadata and trains classification models to identify PHI-related variables.",{"name":82,"@type":73,"acceptedAnswer":83},"How well does the proposed machine learning algorithm perform on unseen datasets?",{"text":84,"@type":76},"Training on multiple large EHR databases from different sources, the algorithm achieves 99% accuracy for detecting PHI-related fields on datasets not seen during training.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]