[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-125418-en":3,"doc-seo-125418-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},125418,1374391974585,"Genevieve","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Exploring machine learning strategies for predicting cardiovascular disease risk factors from multi-omic data","Machine learning classifiers increasingly support prediction of cardiovascular disease (CVD) and related risk factors from omics data, yet the relative influence of the classifier, omic type, and upstream dimension-reduction strategy remains unclear. This study compares six ML classifiers using blood-derived metabolomics, epigenetics, and transcriptomics, with unsupervised or semi-supervised autoencoders for omic reduction. CVD risk factors include blood pressure and ultrasound biomarkers from 1,249 Finnish participants, evaluated with F1 score and transfer learning.","Drouard et al.  \nBMC Medical Informatics and Decision Making [https://doi.org/10.1186/s12911-024-02521-3](https://doi.org/10.1186/s12911-024-02521-3)  \n(2024) 24:116  \nBMC Medical Informatics and Decision Making  \n RESEARCH Open Access  \nExploring machine learning strategies  \nfor predicting cardiovascular disease risk factors from multi-omic data  \nGabin Drouard 1*, Juha Mykkänen2,3, Jarkko Heiskanen2,3, Joona Pohjonen4, Saku Ruohonen3,  \nKatja Pahkala2,3,5, Terho Lehtimäki6, Xiaoling Wang7, Miina Ollikainen1,8, Samuli Ripatti 1,10,9, Matti Pirinen 1,11,9, Olli Raitakari 12,2,3 and Jaakko Kaprio1*  \nAbstract  \nBackground Machine learning (ML) classifiers are increasingly used for predicting cardiovascular disease (CVD) and related risk factors using omics data, although these outcomes often exhibit categorical nature and class imbalances. However, little is known about which ML classifier, omics data, or upstream dimension reduction strategy has the strongest influence on prediction quality in such settings. Our study aimed to illustrate and compare different machine learning strategies to predict CVD risk factors under different scenarios.  \nMethods We compared the use of six ML classifiers in predicting CVD risk factors using blood-derived metabolomics, epigenetics and transcriptomics data. Upstream omic dimension reduction was performed using either unsupervised or semi-supervised autoencoders, whose downstream ML classifier performance we compared. CVD risk factors included systolic and diastolic blood pressure measurements and ultrasound-based biomarkers of left ventricular diastolic dysfunction (LVDD; E/e’ ratio, E/A ratio, LAVI) collected from 1,249 Finnish participants, of which 80% were used for model fitting. We predicted individuals with low, high or average levels of CVD risk factors, the latter class being the most common. We constructed multi-omic predictions using a meta-learner that weighted single-omic predictions. Model performance comparisons were based on the F1 score. Finally, we investigated whether learned omic representations from pre-trained semi-supervised autoencoders could improve outcome prediction in an external cohort using transfer learning.  \nResults Depending on the ML classifier or omic used, the quality of single-omic predictions varied. Multi-omics predictions outperformed single-omics predictions in most cases, particularly in the prediction of individuals with high or low CVD risk factor levels. Semi-supervised autoencoders improved downstream predictions compared to the use of unsupervised autoencoders. In addition, median gains in Area Under the Curve by transfer learning compared to modelling from scratch ranged from 0.09 to 0.14 and 0.07 to 0.11 units for transcriptomic and metabolomic data, respectively.  \n*Correspondence:  \nGabin Drouard [gabin.drouard@helsinki.fi](gabin.drouard@helsinki.fi)[ ](gabin.drouard@helsinki.fi)Jaakko Kaprio[jaakko.kaprio@helsinki.fi](jaakko.kaprio@helsinki.fi)  \nFull list of author information is available at the end of the article  \n© The Author(s) 2024. Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit [http://creativecommons.org/licenses/by/4.0/](http://creativecommons.org/licenses/by/4.0/. The Creative Commons","cbCaieGJoT0uyhSz","https://ap.wps.com/l/cbCaieGJoT0uyhSz","pdf",5436901,1,18,"English","en",105,"# Abstract\n# Background\n# Methods\n## Model building and evaluation\n## Multi-omics and transfer learning\n# Results\n# Conclusions","[{\"question\":\"本研究比较了哪些机器学习策略来预测CVD风险因素？\",\"answer\":\"研究对比了6种机器学习分类器，并在不同情境下评估了使用血源代谢组、表观遗传组与转录组数据时的预测表现，同时比较上游的无监督与半监督自编码器带来的影响。\"},{\"question\":\"CVD风险因素包含哪些指标，以及数据来自谁？\",\"answer\":\"风险因素包括收缩压、舒张压测量，以及基于超声的左室舒张功能障碍相关生物标志物（如E/e’、E/A、LAVI）。数据来自1,249名芬兰参与者，且其中80%用于模型拟合。\"},{\"question\":\"多组学模型与单组学模型、以及迁移学习的效果如何？\",\"answer\":\"多数情况下，多组学预测优于单组学预测，尤其在高/低风险分层方面更明显。半监督自编码器相比无监督自编码器能提升下游预测；迁移学习相对“从头建模”在AUC上带来中位数增益。\"}]","Exploring machine learning strategies for predicting cardiovascular disease risk factors from multi-omic data | PDF",1785898806,45,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"exploring-machine-learning-strategies-for-predicting-cardiovascular-disease-risk-factors-from-multi-omic-data","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/exploring-machine-learning-strategies-for-predicting-cardiovascular-disease-risk-factors-from-multi-omic-data/125418/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"本研究比较了哪些机器学习策略来预测CVD风险因素？","Question",{"text":75,"@type":76},"研究对比了6种机器学习分类器，并在不同情境下评估了使用血源代谢组、表观遗传组与转录组数据时的预测表现，同时比较上游的无监督与半监督自编码器带来的影响。","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"CVD风险因素包含哪些指标，以及数据来自谁？",{"text":80,"@type":76},"风险因素包括收缩压、舒张压测量，以及基于超声的左室舒张功能障碍相关生物标志物（如E/e’、E/A、LAVI）。数据来自1,249名芬兰参与者，且其中80%用于模型拟合。",{"name":82,"@type":73,"acceptedAnswer":83},"多组学模型与单组学模型、以及迁移学习的效果如何？",{"text":84,"@type":76},"多数情况下，多组学预测优于单组学预测，尤其在高/低风险分层方面更明显。半监督自编码器相比无监督自编码器能提升下游预测；迁移学习相对“从头建模”在AUC上带来中位数增益。","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]