[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-126953-en":3,"doc-seo-126953-105":31,"detail-sidebar-cat-0-en-105":84},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},126953,137451207643,"Noah","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","Comparison of Machine Learning Methods for Predicting Homa-IR from Complex Lipid Metabolomics","High-dimensional statistical models are needed to analyze metabolomics data because metabolites are highly correlated and non-normal. Yet metabolomics literature often applies many regression approaches without fully comparing how their underlying differences affect predictive results. This study compares supervised supervised regression methods—MLR, elastic net, PLSR, random forests, and gradient boosting—for predicting HOMA-IR using complex lipid metabolomics data, emphasizing predictive power and identifying key metabolites and covariates across parametric and nonparametric models.","Comparison of Machine Learning Methods for Predicting Homa-IR from Complex Lipid  \nMetabolomics  \nBy Adam Horne  \nSenior Honors Thesis  \nDepartment of Biostatistics  \nUniversity of North Carolina at Chapel Hill  \nApril 2, 2024  \nApproved:  \nDr. Annie Green Howard – Thesis Advisor  \nDr. Katie Meyer – Committee Member  \nJessica Sprinkles – Committee Member  \n1. Abstract  \nHigh-dimensional statistical models are necessary for the analysis of metabolomics data, given the highly correlated nature of the data. However, little research has been done to explore and compare results across commonly used high-dimensional data methods in metabolomics literature. Contemporary studies often use a wide range of regression models without much consideration of their underlying differences. The advent of machine learning has only exacerbated this trend and introduced a plethora of complex, nonlinear methods. However, these complex models may not be appropriate for certain metabolic processes-especially those which are linear in nature. This comparison of selected methods, their predictive power, and the metabolic species highlighted by each aims to explore this issue as it pertains to predicted values of HOMA-IR.  \n2. Introduction  \nThe metabolome is the downstream product of enzymatic reactions in the genome, composed of low molecular weight molecules known as metabolites.1 The measurement of such molecules is the primary interest of metabolomics, which attempts to integrate metabolites into the study and understanding of biological systems. Targeted metabolomics measures known, chemically categorized metabolites and facilitates direct quantitative analysis.2 Predictive models allow researchers to make inferences about biochemical and clinical outcomes,3 but the complexity of targeted metabolomics data complicates statistical processes. Metabolomic profiling yields high-dimensional, non-normal, and often intercorrelated datasets which violate  \nthe assumptions of many traditional models.4 Clinical costs also restrict the sample size of  \ntargeted studies, resulting in “wide” datasets-with fewer subjects relative to the number of variables-that necessitate the use of robust methods.  \nComputational advancements have offset many of these challenges, with growing importance placed on machine learning techniques.5 However, while many of these techniques may prove reliable,6 debates exist over the advantages of machine learning compared to traditional methods.7 This paper will explore the performance of varying supervised regression models frequently used for targeted metabolomics data. The methods of interest are: multiple linear regression (MLR), elastic net regression, partial least squares regression (PLSR), random forests (RF), and gradient boosting. Particular attention will be directed towards the importance of selected metabolites and covariates in each model as well as the relative performance of parametric versus nonparametric methods.  \n3. Methodology  \n3.1. CARDIA Study  \nModel performance was compared using data from the Coronary Artery Risk Development in Young Adults (CARDIA) study, an ongoing prospective cohort designed to investigate risk factors for coronary heart disease.8 From 1985 to 1986, researchers enrolled 5,115 black and white men and women from four U.S. urban population centers: Birmingham, AL; Chicago, IL; Minneapolis, MI; and Oakland, CA.  \nTable 1  \nTable 1 of Year 20 covariates among CLP participants (N=1295) .  \nA targeted complex lipid panel (CLP) was collected in Year 20 and Year 35 and accessed through Metabolon. It measured concentrations of 756 metabolic species across 14 lipid classes. Triacylglycerols (TAGs) comprised 65.3% of tested metabolites with 494 species. A breakdown of the tested metabolites is provided below:  \nTable 2  \nLipid class breakdown of targeted metabolite species (n=756) .  \n\n| Lipid Class (Abbr.) | \\# Tested (% of total) |  |  | Lipid Class (Abbr.) | \\# Tested (% of total) |\n| --- | --- |","cbCaisY38FH1P3mA","https://ap.wps.com/l/cbCaisY38FH1P3mA","pdf",2529317,2,1,22,"English","en",105,"# Abstract\n# Introduction\n# Methodology\n## CARDIA Study\n## HOMA-IR\n## Data Processing","[{\"question\":\"How are missing metabolite concentrations handled?\",\"answer\":\"Metabolites with less than 25% missingness are retained, and missing concentrations are imputed using values drawn from a uniform distribution between the lowest observed value and 1/10th of that lowest observed value for the specific lipid measure.\"}]","Comparison of Machine Learning Methods for Predicting Homa-IR from Complex Lipid Metabolomics | PDF",1785935875,55,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":79,"head_meta":81,"extra_data":83,"updated_unix":29},"comparison-of-machine-learning-methods-for-predicting-homa-ir-from-complex-lipid-metabolomics","",{"@graph":37,"@context":78},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,48,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":20},"https://docshare.wps.com/document/","Document",{"item":49,"name":12,"@type":44,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/comparison-of-machine-learning-methods-for-predicting-homa-ir-from-complex-lipid-metabolomics/126953/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-22","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72],{"name":73,"@type":74,"acceptedAnswer":75},"How are missing metabolite concentrations handled?","Question",{"text":76,"@type":77},"Metabolites with less than 25% missingness are retained, and missing concentrations are imputed using values drawn from a uniform distribution between the lowest observed value and 1/10th of that lowest observed value for the specific lipid measure.","Answer","https://schema.org",{"og:url":52,"og:type":80,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":82,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":85},[86,90,94,98,103,108,113,116,121,124,128],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":87,"show_sort_weight":88,"slug":89},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":91,"show_sort_weight":92,"slug":93},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Exam",70,"exam",{"id":99,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},5,"Comic",60,"comic",{"id":104,"doc_module":4,"doc_module_name":47,"category_name":105,"show_sort_weight":106,"slug":107},6,"Technology",50,"technology",{"id":109,"doc_module":4,"doc_module_name":47,"category_name":110,"show_sort_weight":111,"slug":112},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":114,"slug":115},30,"research-report",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},9,"Religion & Spirituality",20,"religion-spirituality",{"id":119,"doc_module":4,"doc_module_name":47,"category_name":122,"show_sort_weight":119,"slug":123},"World Cup","world-cup",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":125,"slug":127},10,"Lifestyle","lifestyle",{"id":129,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":99,"slug":131},19,"General","general"]