[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120284-en":3,"doc-seo-120284-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120284,13056703019404,"Miles","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Machine Learning Strategies for Improved Phenotype Prediction in Underrepresented Populations","Precision medicine models often deliver better results for European-ancestry populations because genomic datasets and large biobanks over-represent this group. The resulting imbalance can produce less accurate phenotype predictions and treatment recommendations for underrepresented populations, intensifying health disparities. This study presents an adaptable machine learning toolkit that combines existing methods and new techniques to improve prediction accuracy using population-conditional re-sampling. Using UK Biobank SNP data, it yields substantial gains for Asian and African groups, reaching accuracy comparable to the majority.","UC Santa Cruz  \nUC Santa Cruz Previously Published Works  \nTitle  \nMachine Learning Strategies for Improved Phenotype Prediction in Underrepresented Populations.  \nPermalink  \n[https://escholarship.org/uc/item/5xs6t8d5](https://escholarship.org/uc/item/5xs6t8d5)  \nAuthors  \nBonet, David  \nLevin, May Montserrat, Daniel et al.  \nPublication Date  \n2024  \nPeer reviewed  \n[eScholarship.org](eScholarship.org) Powered by the California Digital Library  \nUniversity of California  \nAuthor Manuscr ipt Author Manuscr ipt Author Manuscr ipt Author Manuscript  \n\n|  | HHS Public Access\u003Cbr>Author manuscript\u003Cbr>Pac Symp Biocomput. Author manuscript; available in PMC 2024 January 20. |\n| --- | --- |\n\nPublished in final edited form as: Pac Symp Biocomput. 2024 ; 29: 404–418.  \nMachine Learning Strategies for Improved Phenotype Prediction in Underrepresented Populations  \nDavid Bonet 1,2 , May Levin 1 , Daniel Mas Montserrat 1 , Alexander G. Ioannidis1,3  \n1Stanford University, Stanford, CA, US  \n2 Universitat Politècnica de Catalunya, Barcelona, Spain  \n3 University of California Santa Cruz, Santa Cruz, CA, US  \nAbstract  \nPrecision medicine models often perform better for populations of European ancestry due  \nto the over-representation of this group in the genomic datasets and large-scale biobanks  \nfrom which the models are constructed. As a result, prediction models may misrepresent or provide less accurate treatment recommendations for underrepresented populations, contributing to health disparities. This study introduces an adaptable machine learning toolkit that integrates multiple existing methodologies and novel techniques to enhance the prediction accuracy for underrepresented populations in genomic datasets. By leveraging machine learning techniques, including gradient boosting and automated methods, coupled with novel population-conditional re-sampling techniques, our method significantly improves the phenotypic prediction from single nucleotide polymorphism (SNP) data for diverse populations. We evaluate our approach using the UK Biobank, which is composed primarily of British individuals with European ancestry, and a minority representation of groups with Asian and African ancestry. Performance metrics demonstrate substantial improvements in phenotype prediction for underrepresented groups, achieving prediction accuracy comparable to that of the majority group. This approach represents a significant step towards improving prediction accuracy amidst current dataset diversity challenges. By integrating a tailored pipeline, our approach fosters more equitable validity and utility of statistical genetics methods, paving the way for more inclusive models and outcomes.  \nKeywords  \nGenetics; Precision Medicine; Machine Learning; Phenotype Prediction; Bioinformatics  \n1. Introduction  \nIn recent years, genome-wide association studies (GWAS) have provided many insights into the genetic basis of complex traits and diseases. However, these findings predominantly benefit populations of European descent due to their over-representation in genomic  \nOpen Access chapter published by World Scientific Publishing Company and distributed under the terms of the Creative Commons Attribution Non-Commercial (CC BY-NC) 4.0 License.  \n[ioannidis@stanford.edu](ioannidis@stanford.edu) .  \nAuthor Manuscr ipt Author Manuscr ipt Author Manuscr ipt Author Manuscript  \nBonet et al. Page 2  \ndatasets. Individuals with Asian, African, and other ancestries only represent a small fraction of the available datasets.1 Although individuals of European descent constitute ∼79% of GWAS participants,2 they account for less than a quarter of the global population. This disproportionate representation creates a limitation in precision medicine, because statistical models built to infer disease risks or health-related traits can perform poorly for individuals from populations that were underrepresented when creating the model, exacerbating health disparities. Despite initia","cbCaifVTzidrBKX6","https://ap.wps.com/l/cbCaifVTzidrBKX6","pdf",1142073,1,20,"English","en",105,"# Introduction\n## Background and dataset imbalance\n## Phenotype prediction and modeling approaches\n# Bias in ML-based prediction pipelines\n## Limitations of naive ML applications\n## Existing bias mitigation methods","[{\"question\":\"Why do phenotype prediction models often underperform for underrepresented populations?\",\"answer\":\"Genomic datasets and biobanks over-represent European ancestry, so models trained on these data can misrepresent or predict less accurately for populations that were underrepresented.\"},{\"question\":\"What does the study introduce to address prediction bias?\",\"answer\":\"An adaptable machine learning toolkit that integrates multiple methodologies and novel population-conditional re-sampling techniques to enhance phenotype prediction accuracy.\"},{\"question\":\"How is the proposed approach evaluated and what is the outcome?\",\"answer\":\"The method is evaluated on UK Biobank SNP data, where performance metrics show substantial improvements for underrepresented groups, achieving prediction accuracy comparable to the majority group.\"}]","Machine Learning Strategies for Improved Phenotype Prediction in Underrepresented Populations | PDF",1785729237,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"machine-learning-strategies-for-improved-phenotype-prediction-in-underrepresented-populations","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/machine-learning-strategies-for-improved-phenotype-prediction-in-underrepresented-populations/120284/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do phenotype prediction models often underperform for underrepresented populations?","Question",{"text":75,"@type":76},"Genomic datasets and biobanks over-represent European ancestry, so models trained on these data can misrepresent or predict less accurately for populations that were underrepresented.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does the study introduce to address prediction bias?",{"text":80,"@type":76},"An adaptable machine learning toolkit that integrates multiple methodologies and novel population-conditional re-sampling techniques to enhance phenotype prediction accuracy.",{"name":82,"@type":73,"acceptedAnswer":83},"How is the proposed approach evaluated and what is the outcome?",{"text":84,"@type":76},"The method is evaluated on UK Biobank SNP data, where performance metrics show substantial improvements for underrepresented groups, achieving prediction accuracy comparable to the majority group.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":29,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":21,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":21,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]