[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117640-en":3,"doc-seo-117640-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117640,962075114101,"Seraphina","https://ap-avatar.wpscdn.com/avatar/e000253a75eb197efd?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780044092746381165",8,"Research & Report","Machine learning methods for detecting positive selection","Molecular evolutionary biology aims to explain the diversity of life by identifying genomic processes that shape genomes over time, with adaptation often revealed as signatures of positive selection on protein-coding genes. Reliable detection of these genomic footprints remains difficult, as classical likelihood-based approaches for interspecific positive selection rely on codon substitution models and codon-rate inference from dN/dS. These methods can be computationally demanding, sensitive to model misspecification, and prone to errors from simplifying assumptions and misalignment. This thesis develops machine learning methods—including CNN and transformer models—trained on simulated data to classify positive selection, evaluate benchmarking against likelihood methods, and extend to sitewise dN/dS prediction with attention to interpretability.","Machine learning methods for detecting positive selection  \nCharlotte Edra Mary West  \nEMBL-European Bioinformatics Institute Darwin College, University of Cambridge October, 2025  \nThis thesis is submitted for the degree of Doctor of Philosophy  \nDeclaration  \nThis thesis is the result of my own work and includes nothing which is the outcome of work done in collaboration except as declared in the preface and specified in the text. It is not substantially the same as any work that has already been submitted, or is being concurrently submitted, for any degree, diploma or other qualification at the University of Cambridge or any other University or similar institution except as declared in the preface and specified in the text. It does not exceed the prescribed word limit for the relevant Degree Committee.  \nCharlotte Edra Mary West October, 2025  \nAbstract  \nMachine learning methods for detecting positive selection  \nCharlotte Edra Mary West  \nMolecular evolutionary biology seeks to explain the diversity of life by uncovering the processes that shape genomes over time. A central aim within this field is to identify the genetic basis of adaptation, often manifesting as signatures of positive selection acting on protein-coding genes. Detecting such signals sheds light on evolutionary processes such as functional divergence, coevolutionary dynamics and phenotypic innovation. However, whilst the study of natural selection has been foundational in evolutionary theory, reliably identifying its genomic footprints remains a persistent challenge.  \nTraditional methods for detecting interspecific positive selection are grounded in statistical, likelihood-based methods, typically employing codon substitution models. These approaches infer rates of nonsynonymous to synonymous substitutions (dN/dS) from nucleotide multiple sequence alignments (MSAs) of homologous, protein-coding genes, interpreted as a proxy for positive selection. These approaches have been instrumental in advancing our understanding of adaptive evolution, but they are not without their limitations. Statistical approaches are often computationally demanding, vulnerable to model misspecification, and underpowered when applied to realistic evolutionary scenarios. Models are forced to make simplifying assumptions to make the analysis tractable in the likelihood framework, which can contribute to both type I and type II error. Misalignments tend to further inflate false positive inference. As genomic datasets have grown in both scale and complexity, the shortcomings of classical inference methods have become increasingly apparent. This creates a pressing need for approaches that can harness large genomic datasets more flexibly, while capturing the complex dependencies inherent in sequence evolution.  \nIn parallel with these challenges, recent advances in machine learning have transformed diverse areas of biology, including molecular evolution. In particular, deep neural network architectures such as convolutional neural networks (CNNs) and transformer models have demonstrated an ability to extract meaningful patterns from raw biological data, without heavy reliance on human-curated features. Nevertheless, applying machine learning  \nto molecular evolution presents its own challenges: the task requires large amounts of biologically realistic training data, architectures and resources equipped to handle such data, and rigorous evaluation frameworks that connect machine learning predictions to established evolutionary theory. There is very little or no evolutionary data for which we know the ground truth regarding how it has evolved; therefore, I rely on simulated data for training and evaluation.  \nThis thesis addresses these challenges by developing and evaluating machine learning methods for detecting positive selection in protein-coding genes. I first establish simulation frameworks that can be used to generate training data, and design benchmarking experiments to compare mac","cbCaiuE8wg8qm38T","https://ap.wps.com/l/cbCaiuE8wg8qm38T","pdf",32051364,1,186,"English","en",105,"# Abstract\n## Motivation and challenge\n## Classical likelihood-based detection and limitations\n## Machine learning opportunity and constraints\n## Thesis approach: CNNs\n## Transformer models and sitewise inference\n## Evaluation and interpretation","[{\"question\":\"What biological problem does the thesis focus on?\",\"answer\":\"The thesis focuses on detecting positive selection in protein-coding genes by identifying genomic signatures of adaptation over evolutionary time.\"},{\"question\":\"Why are classical interspecific positive selection methods considered challenging?\",\"answer\":\"They can be computationally demanding, vulnerable to model misspecification, underpowered in realistic scenarios, and affected by simplifying assumptions and misalignments that can increase false inferences.\"},{\"question\":\"How does the thesis use machine learning to detect positive selection?\",\"answer\":\"It develops and evaluates CNN and transformer models trained on simulated data to classify MSAs under positive selection or not, and transformer models further predict sitewise dN/dS values for codon-specific inference.\"}]","Machine learning methods for detecting positive selection | PDF",1785677553,469,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"machine-learning-methods-for-detecting-positive-selection","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/machine-learning-methods-for-detecting-positive-selection/117640/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What biological problem does the thesis focus on?","Question",{"text":75,"@type":76},"The thesis focuses on detecting positive selection in protein-coding genes by identifying genomic signatures of adaptation over evolutionary time.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why are classical interspecific positive selection methods considered challenging?",{"text":80,"@type":76},"They can be computationally demanding, vulnerable to model misspecification, underpowered in realistic scenarios, and affected by simplifying assumptions and misalignments that can increase false inferences.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the thesis use machine learning to detect positive selection?",{"text":84,"@type":76},"It develops and evaluates CNN and transformer models trained on simulated data to classify MSAs under positive selection or not, and transformer models further predict sitewise dN/dS values for codon-specific inference.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]