[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119599-en":3,"doc-seo-119599-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119599,687197207919,"Theodora","https://ap-avatar.wpscdn.com/avatar/a000253d6f5f7c60be?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779446848396160552",8,"Research & Report","Machine learning in Alzheimer’s disease genetics","Machine learning models are applied to genome-wide data from 41,686 individuals in a large European Alzheimer’s disease consortium to evaluate algorithm performance, replicate established genetic findings, discover novel loci, and predict individual risk. Gradient Boosting Machines, biological pathway-informed neural networks, and model-based multifactor dimensionality reduction are compared. The models recover genome-wide significant variants and capture a substantial share of meta-analysis associations, identifying replicated novel loci and refining the SPPL2A locus. Results show predictive performance comparable to classical methods and highlight ML’s potential to uncover signals missed by traditional GWAS.","Article [https://doi.org/10.1038/s41467-025-61650-z](https://doi.org/10.1038/s41467-025-61650-z)  \nMachine learning in Alzheimer’s disease genetics  \nReceived: 26 July 2024  \n\n| Accepted: 24 June 2025 |\n| --- |\n|  |\n| Check for updates |\n\nA list of authors and their afﬁliations appears at the end of the paper  \nTraditional statistical approaches have advanced our understanding of the genetics of complex diseases, yet are limited to linear additive models. Here we applied machine learning (ML) to genome-wide data from 41,686 individuals in the largest European consortium on Alzheimer’s disease (AD) to investigate the effectiveness of various ML algorithms in replicating known ﬁndings, discovering novel loci, and predicting individuals at risk. We utilised Gradient Boosting Machines (GBMs), biological pathway-informed Neural Networks (NNs), and Model-based Multifactor Dimensionality Reduction (MBMDR) models. ML approaches successfully captured all genome-wide signiﬁcant genetic variants identiﬁed in the training set and 22% of associations from larger meta-analyses. They highlight 6 novel loci which replicate in an external dataset, including variants which map to ARHGAP25, LY6H, COG7, SOD1 and ZNF597. They further identify novel association in AP4E1, reﬁning the genetic landscape of the known SPPL2A locus. Our results demonstrate that machine learning methods can achieve predictive performance comparable to classical approaches in genetic epidemiology and have the potential to uncover novel loci that remain undetected by traditional GWAS. These insights provide a complementary avenue for advancing the understanding of AD genetics.  \nGenome-wide association studies (GWAS) have enabled huge progress in identifying variants associated with the risk of developing Alzheimer’s disease (AD)1. Polygenic risk scores (PRS) based on these variants have greatly improved prediction ofdisease status2. However, inherent to GWAS and PRS are the assumptions that variants are independent predictors, linearly associated with the outcome, and therefore combine additively within and between loci3, with no interactions occurring between variants, or between genes and other risk factors. While such simplifying genetic assumptions have proved fruitful across a range of diseases and disorders4,5, they are at odds with biological evidence in AD that disease heterogeneity and responses from cells such as microglia are dependent on APOE status6–9. Further, there is genetic evidence suggesting that different variants are associated with the disease depending on APOE status10–13 and age at diagnosis or assessment14–16. As GWAS sample size increases and PRS approach limits on predictive performance, alternative modelling approaches  \nare essential to maximise discoveries from existing data and enable a deeper understanding of AD genetics.  \nThe conﬂuence of increasingly large genetic data17, readily available computational resources, and mature methodologies presents a key opportunity for addressing this at scale by applying ﬂexible datadriven machine learning (ML) models. Several studies have applied ML to the genetics of brain disorders and have been recently summarised18,19. Previous ML attempts have been impacted by high risk of bias20 and population stratiﬁcation21, while AD studies in particular have been hampered by low sample size18, leaving a gap for comprehensive, large-scale studies that rigorously apply ML to genome-wide data. We conducted the largest genome-wide ML study in AD to date, marking a pivotal moment in the ﬁeld. Our study presents a reproducible, bias-aware approach for ML model development, validation, and confounder adjustment. We trained three of the most prominent approaches in the ﬁeld to compare predictive accuracy and uncover  \ne-mail: [kristel.vansteen@uliege.be](kristel.vansteen@uliege.be); [cornelia.vanduijn@ndph.ox.ac.uk](cornelia.vanduijn@ndph.ox.ac.uk); [EscottPriceV@cardiff.ac.uk](EscottPriceV@cardiff.ac.uk)  \nFig. 1 | M","cbCaimr5zt1ULUtO","https://ap.wps.com/l/cbCaimr5zt1ULUtO","pdf",3588285,1,16,"English","en",105,"# Methods overview\n## Data split, cross-validation, and hyperparameter tuning\n## Association analysis: annotation, enrichment, interaction testing, replication\n## Prediction evaluation: AUC and correlations\n# Results: Prediction performance comparison\n## Correlations across models and repeats\n## Case-control discrimination using AUC","[{\"question\":\"Which machine learning algorithms were evaluated for Alzheimer’s disease genetics?\",\"answer\":\"Gradient Boosting Machines, biological pathway-informed neural networks, and model-based multifactor dimensionality reduction were evaluated, alongside polygenic risk score models.\"},{\"question\":\"How well did the models replicate known genome-wide genetic findings?\",\"answer\":\"They captured all genome-wide significant genetic variants in the training set and reproduced 22% of associations reported in larger meta-analyses.\"},{\"question\":\"What predictive performance did the models achieve compared with classical approaches?\",\"answer\":\"Gradient boosting achieved the highest discrimination with an AUC of 0.692, which was not significantly different from the AUC of 0.689 for polygenic risk scores, and performance remained stable across splits and cohorts.\"}]","Machine learning in Alzheimer’s disease genetics | PDF",1785725215,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"machine-learning-in-alzheimers-disease-genetics","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/machine-learning-in-alzheimers-disease-genetics/119599/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Which machine learning algorithms were evaluated for Alzheimer’s disease genetics?","Question",{"text":75,"@type":76},"Gradient Boosting Machines, biological pathway-informed neural networks, and model-based multifactor dimensionality reduction were evaluated, alongside polygenic risk score models.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How well did the models replicate known genome-wide genetic findings?",{"text":80,"@type":76},"They captured all genome-wide significant genetic variants in the training set and reproduced 22% of associations reported in larger meta-analyses.",{"name":82,"@type":73,"acceptedAnswer":83},"What predictive performance did the models achieve compared with classical approaches?",{"text":84,"@type":76},"Gradient boosting achieved the highest discrimination with an AUC of 0.692, which was not significantly different from the AUC of 0.689 for polygenic risk scores, and performance remained stable across splits and cohorts.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":29,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]