[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118786-en":3,"doc-seo-118786-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118786,549758146520,"Patrick","https://ap-avatar.wpscdn.com/avatar/80002397d8c0411e94?_k=1775819394049821470",8,"Research & Report","Omics Data Preprocessing for Machine Learning - A Case Study in Childhood Obesity","Machine learning is increasingly used to build predictive models for disease outcomes using omics and other molecular data, but success depends on rigorous algorithm use and, critically, correct preprocessing and management of input datasets. Many omics-and-ML approaches fail across key steps including experimental design, feature selection, data preprocessing, and algorithm choice. This work provides a practical guideline and best-practice recommendations for addressing major challenges in multi-omics human studies, such as biological heterogeneity, technical noise, high dimensionality, missing values, and class imbalance.","Article  \nOmics Data Preprocessing for Machine Learning: A Case Study in Childhood Obesity  \nÁlvaro Torres-Martos 1,2,3, Mireia Bustos-Aibar 1,2,3, Alberto Ramírez-Mena 4, Sofía Cámara-Sánchez 5, Augusto Anguita-Ruiz 2,3,6,7, *, Rafael Alcalá 5, Concepción M. Aguilera 1,2,3,7 and Jesús Alcalá-Fdez 5  \nCitation: Torres-Martos, Á.;  \nBustos-Aibar, M.; Ramírez-Mena, A.; Cámara-Sánchez, S.; Anguita-Ruiz, A.; Alcalá, R.; Aguilera, C.M.; Alcalá-Fdez, J. Omics Data Preprocessing for Machine Learning:  \nA Case Study in Childhood Obesity. Genes 2023, 14, 248. [https://doi.org/](https://doi.org/)[ ](https://doi.org/)[10.3390/genes14020248](10.3390/genes14020248)  \nAcademic Editors: Francisco Ortuño, Alfredo Benso, Jean-Marc Schwartz, Alexandre G. de Brevern, Ignacio Rojas and Olga Valenzuela  \nReceived: 15 October 2022  \nRevised: 11 January 2023  \nAccepted: 12 January 2023  \nPublished: 18 January 2023  \nCopyright: © 2023 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license ([https://](https://)[ ](https://)[creativecommons.org/licenses/by/](creativecommons.org/licenses/by/)[ ](creativecommons.org/licenses/by/)[4.0/](4.0/)) .  \n1 Department of Biochemistry and Molecular Biology II, University of Granada, 18071 Granada, Spain  \n2 “José Mataix Verdú” Institute of Nutrition and Food Technology (INYTA), Center of Biomedical Research, University of Granada, 18100 Granada, Spain  \n3 Biosanitary Research Institute of Granada (IBS.GRANADA), 18012 Granada, Spain  \n4 Centre for Genomics and Oncological Research (GENYO), 18016 Granada, Spain  \n5 Department of Computer Science and Artiﬁcial Intelligence, Andalusian Research Institute in Data Science and Computational Intelligence (DaSCI), University of Granada, 18071 Granada, Spain  \n6 Barcelona Institute for Global Health (ISGlobal), 08003 Barcelona, Spain  \n7 CIBER Physiopathology of Obesity and Nutrition (CIBEROBN), Instituto de Salud Carlos III,  \n28029 Madrid, Spain  \n* Correspondence: [augusto.anguita@isglobal.org](augusto.anguita@isglobal.org)  \n† These authors contributed equally to this work.  \nAbstract: The use of machine learning techniques for the construction of predictive models of disease outcomes (based on omics and other types of molecular data) has gained enormous relevance in the last few years in the biomedical ﬁeld. Nonetheless, the virtuosity of omics studies and machine learning tools are subject to the proper application of algorithms as well as the appropriate preprocessing and management of input omics and molecular data. Currently, many of the available approaches that use machine learning on omics data for predictive purposes make mistakes in several of the following key steps: experimental design, feature selection, data pre-processing, and algorithm selection. For this reason, we propose the current work as a guideline on how to confront the main challenges inherent to multi-omics human data. As such, a series of best practicesand recommendations are also presented for each of the steps deﬁned. In particular, the main particularities of each omics data layer, the most suitable preprocessing approaches for each source, and a compilation of best practices and tips for the study of disease development prediction using machine learning are described. Using examples of real data, we show how to address the key problems mentioned in multi-omics research (e.g., biological heterogeneity, technical noise, high dimensionality, presence of missing values, and class imbalance) . Finally, we deﬁne the proposals for model improvement based on the results found, which serve as the bases for future work.  \nKeywords: machine learning; omics; data pre-processing  \n1. Introduction  \nIn recent years, the biomedical ﬁeld has experienced a big data revolution. Since the appearance of the ﬁrst microarray technologies, the competencies of generating data and extra","cbCaicDTgDL5omSN","https://ap.wps.com/l/cbCaicDTgDL5omSN","pdf",1557322,1,16,"English","en",105,"# Introduction\n## Omics data and predictive modeling\n## Multi-omics challenges and preprocessing needs","[{\"question\":\"What is the main focus of the proposed work on multi-omics machine learning?\",\"answer\":\"It provides a guideline addressing the main challenges in multi-omics human data and outlines best practices for each major step, especially preprocessing and data management.\"},{\"question\":\"Which key steps are identified as common sources of mistakes in omics-based predictive modeling?\",\"answer\":\"The document highlights experimental design, feature selection, data pre-processing, and algorithm selection as frequent problem areas.\"},{\"question\":\"What preprocessing challenges are discussed using examples of real data?\",\"answer\":\"Examples address biological heterogeneity, technical noise, high dimensionality, missing values, and class imbalance.\"}]","Omics Data Preprocessing for Machine Learning - A Case Study in Childhood Obesity | PDF",1785720245,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"omics-data-preprocessing-for-machine-learning-a-case-study-in-childhood-obesity","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/omics-data-preprocessing-for-machine-learning-a-case-study-in-childhood-obesity/118786/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the main focus of the proposed work on multi-omics machine learning?","Question",{"text":75,"@type":76},"It provides a guideline addressing the main challenges in multi-omics human data and outlines best practices for each major step, especially preprocessing and data management.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which key steps are identified as common sources of mistakes in omics-based predictive modeling?",{"text":80,"@type":76},"The document highlights experimental design, feature selection, data pre-processing, and algorithm selection as frequent problem areas.",{"name":82,"@type":73,"acceptedAnswer":83},"What preprocessing challenges are discussed using examples of real data?",{"text":84,"@type":76},"Examples address biological heterogeneity, technical noise, high dimensionality, missing values, and class imbalance.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":29,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]