[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-124345-en":3,"doc-seo-124345-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},124345,962075114101,"Seraphina","https://ap-avatar.wpscdn.com/avatar/e000253a75eb197efd?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780044092746381165",8,"Research & Report","Machine Learning Applications in Epigenomics - its Association with Health and Disease - thesis for the degree of Doctor of Philosophy","Epigenetic modifications alter phenotypes without changing the DNA sequence, and disrupted epigenetically regulated gene expression contributes to cancers, autoimmune diseases, and other disorders. Machine learning (ML) leverages trained algorithms to extract patterns from data, and is increasingly used in epigenomic studies such as DNA methylation (DNAm) analysis, where high dimensionality, computational burden, and overfitting make feature reduction essential. This thesis evaluates robust, systematic ML workflows to build DNAm-based predictors for telomere length, oesophageal cancer, and inflammatory bowel disease, identifying optimal models and integrating the methods into a public package.","Technological University Dublin  \nMachine Learning Applications in Epigenomicsand its Association with Health and Disease  \nTrevor Doherty, MSc, PgDip (Res) , PgDip, BSc  \nA thesis submitted for the degree of Doctor of Philosophy  \nUnder the supervision of  \nDr. Therese Murphy & Prof. Sarah Jane Delany  \nSchool of Biological, Health and Sports Sciences  \nNovember 2024  \nAbstract  \nEpigenetic modifications can lead to altered phenotypes without a change in the DNA sequence itself. Disrupted gene expression regulated by epigenetic processes can result in cancers, autoimmune diseases and various other maladies. Machine learning (ML) involves the use of algorithms and models which are trained to learn patterns in data, and has demonstrated remarkable success in solving diverse, complex challenges. Epigenomic studies, such as those that use DNA methylation (DNAm) data, increasingly make use of ML techniques to process extremely high dimensional data obtained from highthroughput platforms e.g., DNAm arrays. These datasets suffer from the curse of dimensionality, increased computational complexity and are prone to overfitting–making feature reduction techniques critical.  \nMany DNAm-based studies frequently test a single or small range of feature reduction approaches and ML algorithms when developing prediction models. Moreover, researchers do not necessarily choose a feature reduction or prediction algorithm that is appropriate for their data. In contrast, in the ML domain, multiple algorithms are typically evaluated due to dataset dependence. Therefore, the research in this thesis evaluates the application of ML, using robust systematic data-driven practices and methodologies, in the development of novel DNAm-based prediction models of health and disease, finding the optimal model in three DNAm-based prediction tasks-telomere length estimation, oesophageal cancer (OC) classification and inflammatory bowel diseases (IBD) prediction.  \nIn the case of the telomere length estimators, we applied the methodology to develop multiple novel DNAm-based signatures of aging. Using diverse feature reduction approaches, we showed substantial variation in the models’ ability to predict – thus  \nhighlighting the critical need for a systematic approach. The best estimator identified, PCA-EN TL, used principal component analysis in advance of Elastic Net regression.  \nBeyond predicting quantitative traits, optimal novel classification models of OC and IBD were developed by evaluating a broad array of feature reduction and ML approaches. OC is one of the leading causes of cancer-related deaths and is often asymptomatic in its early stages. Frequent late detection leads to poor prognosis, making timely intervention using non-invasive biomarkers critical. The optimal novel DNAmbased OC classifier showed strong performance in discriminating normal, Barrett’soesophagus (BO) and OC, achieving sensitivities of 95.7%, 92.4% and 82.1% respectively with a balanced accuracy of 90.1% . Novel few-CpG signatures have the potential to enable cost-effective DNAm assays and could be tested in liquid biopsies. As such, we identified a 27-CpG signature, which achieved a balanced accuracy of 87.6% and normal, BO and OC sensitivities of 91.5%, 75% and 96.3% respectively. Moreover, the signature demonstrated potential to stratify early and advanced oesophageal squamous cell carcinoma.  \nRegarding IBD, we examined multiple second- and third-generation epigenetic clocks and aging signatures, finding significant associations with IBD outcomes in both discovery and replication cohorts. Additionally, we identified evidence for DunedinPACE as a more effective biomarker of inflammation than C-reactive protein in ulcerative colitis patients. Finally, our robust systematic classification methodology was integrated into a publicly available package which can be used by other researchers who wish to develop their own optimal novel DNAm-based classifiers of health- and disease-relat","cbCaigc3qQvUBkWk","https://ap.wps.com/l/cbCaigc3qQvUBkWk","pdf",11185698,1,425,"English","en",105,"# Abstract\n## ML and epigenomic prediction challenges\n## Evaluation of DNAm-based prediction tasks\n## Telomere length estimation\n## Oesophageal cancer classification\n## Inflammatory bowel disease prediction\n## Public integration of the methodology","[{\"question\":\"Why are feature reduction techniques critical for DNAm-based machine learning models?\",\"answer\":\"DNAm datasets are high dimensional, computationally complex, and prone to overfitting. Feature reduction helps address dimensionality and improves modeling robustness.\"},{\"question\":\"Which method was highlighted as the best telomere length estimator in the thesis?\",\"answer\":\"The best estimator identified is PCA-EN TL, which applies principal component analysis before Elastic Net regression.\"},{\"question\":\"How does the thesis support the use of DNAm biomarkers for oesophageal cancer?\",\"answer\":\"It develops and evaluates multiple feature reduction and ML approaches, producing classifiers that distinguish normal, Barrett’s oesophagus, and oesophageal cancer with strong sensitivity and balanced accuracy. It also identifies a 27-CpG signature with performance suitable for cost-effective DNAm assays and potential liquid biopsy testing.\"}]","Machine Learning Applications in Epigenomics - its Association with Health and Disease - thesis for the degree of Doctor of Philosophy | PDF",1785821750,1071,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"machine-learning-applications-in-epigenomics-its-association-with-health-and-disease-thesis-for-the-degree-of-doctor-of-philosophy","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/machine-learning-applications-in-epigenomics-its-association-with-health-and-disease-thesis-for-the-degree-of-doctor-of-philosophy/124345/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why are feature reduction techniques critical for DNAm-based machine learning models?","Question",{"text":75,"@type":76},"DNAm datasets are high dimensional, computationally complex, and prone to overfitting. Feature reduction helps address dimensionality and improves modeling robustness.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which method was highlighted as the best telomere length estimator in the thesis?",{"text":80,"@type":76},"The best estimator identified is PCA-EN TL, which applies principal component analysis before Elastic Net regression.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the thesis support the use of DNAm biomarkers for oesophageal cancer?",{"text":84,"@type":76},"It develops and evaluates multiple feature reduction and ML approaches, producing classifiers that distinguish normal, Barrett’s oesophagus, and oesophageal cancer with strong sensitivity and balanced accuracy. It also identifies a 27-CpG signature with performance suitable for cost-effective DNAm assays and potential liquid biopsy testing.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]