[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120452-en":3,"doc-seo-120452-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120452,962075114765,"Quinn","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","From EHR to Machine Learning - A Preliminary Report on an Ingestion Pipeline Based on JSON-LD","The paper presents preliminary experiments for building an ingestion mechanism that transfers data from Electronic Health Records (EHR) to machine learning pipelines using Linked Data and the JSON-LD format. It motivates the approach by highlighting interoperability gaps and the limitations of converting coded EHR data into simpler formats such as CSV. Using an Italian epidemiological information system, the study structures coded data linkage through a JSON-LD context to preserve provenance and support model training and inference.","818  \nDigital Health and Informatics Innovations for Sustainable Health Care Systems  \nJ. Mantas et al. (Eds.)© 2024 The Authors.  \nThis article is published online with Open Access by IOS Press and distributed under the terms of the Creative Commons Attribution Non-Commercial License 4.0 (CC BY-NC 4.0).  \ndoi:10.3233/SHTI240536  \nFrom EHR to Machine Learning: A Preliminary Report on an Ingestion Pipeline Based on JSON-LD  \nGiulia Lucrezia BARONIa , Vincenzo DELLA MEA a,1 and Gian Luca FORESTIa aDept. of Mathematics, Computer Science and Physics, University of Udine, Italy ORCiD ID: Vincenzo Della Mea [https://orcid.org/0000-0002-0144-3802](https://orcid.org/0000-0002-0144-3802)  \nAbstract. In this paper, we present the preliminary experiments for the development of an ingestion mechanism to move data from Electronic Health Records to machine learning processes, based on the concept of Linked Data and the JSON-LD format.  \nKeywords. EHR, Machine Learning, Linked Data  \n1. Introduction  \nIn the last few years, the availability of novel machine learning methods fostered advances in various areas of clinical practice, starting from those involving biomagesand biosignals. However, of particular interest is also machine learning applied to health data extracted from Electronic Health Records (EHR), at least is sectors where data collection happens mostly in electronic form. EHRs are commonly supported by relational databases. Furthermore, part of the data can be coded using terminologies and classifications like ICD, SNOMED-CT, LOINC, and other. Finally, sometimes EHRscan communicate with other health information systems by means of the available health informatics standards, like HL7, CDA and FHIR. When experimenting and implementing machine learning (ML) within health information systems, the first step is to build a preprocessing pipeline able to extract the needed data from the EHR in a suitable form for model ingestion. We report on preliminary experiments carried out in the framework ofa Rare Diseases project.  \n2. Methods  \nWe based our analysis on the Epidemiological Information System of the Region Friuli Venezia Giulia, Italy. By examining the underlying database, consisting of 23 modules for a total of 179 tables, we found that there were at least 9 dictionaries potentially related to internationally defined artifacts (diseases, procedures, drugs, laboratory exams, etc) and other 11 locally defined (geographical locations, health structures, etc) .  \nNotwithstanding the abundant data currently digitized in EHRs, recent evidence demonstrates that only a fraction of such data – on average 27 variables, and mostly from  \n1 Corresponding Author: Vincenzo Della Mea, [E-mail: vincenzo.dellamea@uniud.it](E-mail: vincenzo.dellamea@uniud.it).  \nG.L. Baroni et al. / From EHR to Machine Learning: A Preliminary Report 819  \na single system-is commonly exploited in machine learning [2] . One reason is the lack of technical or semantic interoperability. One common format is CSV, being it easily manageable with all machine learning frameworks and libraries, but it lacks the capability of fully representing the richness of coded data. Having this in mind, we wanted to identify a possible path to simplify the ingestion of data, taking also in account the richness given by dictionaries, terminologies and classifications, which is often lost when data is converted to CSV or equivalent formats.  \n3. Results  \nOur proposal relies on the Linked Data concept, and in particular on the JSON-LD format. JSON-LD provides a lightweight framework for linked data, that is, data interconnected with other data. In our case, EHR coded data should be connected with biomedical dictionaries in an explicit way, otherwise lost with simpler formats like CSV. JSON-LD is centered around the ‘context’, which allows to relate one or more variables to concepts specified in ontologies. In our case, it is straightforward to use this approach to link coded data with t","cbCaia8uhZYFWN31","https://ap.wps.com/l/cbCaia8uhZYFWN31","pdf",158450,1,2,"English","en",105,"# Introduction\n# Methods\n# Results\n# Discussion","[{\"question\":\"What problem does the paper address when moving from EHR to machine learning?\",\"answer\":\"It addresses the need for a preprocessing/ingestion pipeline that extracts EHR data in a model-ingestion-friendly form while preserving the richness of coded information and overcoming interoperability limitations.\"},{\"question\":\"Why is JSON-LD used in the proposed ingestion pipeline?\",\"answer\":\"JSON-LD supports Linked Data by using a context to explicitly relate EHR variables to concepts in biomedical dictionaries/ontologies, enabling provenance tracking during training and inference.\"},{\"question\":\"What data source and system are used for the preliminary experiments?\",\"answer\":\"The experiments are based on the Epidemiological Information System of the Region Friuli Venezia Giulia, Italy, examining a database with multiple modules, tables, and dictionaries for both internationally defined and locally defined artifacts.\"}]","From EHR to Machine Learning - A Preliminary Report on an Ingestion Pipeline Based on JSON-LD | PDF",1785730181,5,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"from-ehr-to-machine-learning-a-preliminary-report-on-an-ingestion-pipeline-based-on-json-ld","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":21},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/from-ehr-to-machine-learning-a-preliminary-report-on-an-ingestion-pipeline-based-on-json-ld/120452/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-08-04","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address when moving from EHR to machine learning?","Question",{"text":75,"@type":76},"It addresses the need for a preprocessing/ingestion pipeline that extracts EHR data in a model-ingestion-friendly form while preserving the richness of coded information and overcoming interoperability limitations.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why is JSON-LD used in the proposed ingestion pipeline?",{"text":80,"@type":76},"JSON-LD supports Linked Data by using a context to explicitly relate EHR variables to concepts in biomedical dictionaries/ontologies, enabling provenance tracking during training and inference.",{"name":82,"@type":73,"acceptedAnswer":83},"What data source and system are used for the preliminary experiments?",{"text":84,"@type":76},"The experiments are based on the Epidemiological Information System of the Region Friuli Venezia Giulia, Italy, examining a database with multiple modules, tables, and dictionaries for both internationally defined and locally defined artifacts.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":29,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":29,"slug":137},19,"General","general"]