[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119972-en":3,"doc-seo-119972-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119972,4398048949847,"Eliana","https://ap-avatar.wpscdn.com/avatar/400002536579ef2da7f?_k=1778318612642679267",8,"Research & Report","Machine Learning for De Novo Peptide Identification - Doctoral Thesis","Proteomics supports ecosystem-level biological insight by identifying and analyzing proteins, traditionally by matching peptide subsequences from tandem mass spectra via database searching. Yet a substantial fraction of experimental spectra—reported on average as 75%—remain unidentified. De novo peptide identification addresses this by inferring peptide sequences directly from spectra and has recently improved through machine learning–enhanced algorithms. This thesis evaluates strengths and weaknesses of current state-of-the-art de novo methods, surveys tandem mass spectral characteristics, proposes an alternative CNN-GNN ion encoding module, and examines how artificial spectra can be noise-augmented to better mimic real data for training and testing.","School of Computer Science College of Science and Engineering University of Galway  \nMachine Learning for De Novo Peptide  \nIdentification  \nKevin McDonnell  \nSupervised by:  \nDr. Enda Howley and Dr. Florence Abram  \nA thesis submitted for the degree of Doctor of Philosophy  \nJanuary 2023  \nAbstract  \nProteomics involves the identification and analysis of proteins, therefore providing valuable insight into ecosystem functioning. In this methodology, protein sequences are typically identified using a bottom-up approach whereby short subsequences called peptides are matched to experimental mass spectra using a database search. However, it is reported that on average, 75% of the spectra recovered from experiments remain unidentified. De novo peptide identification is an alternative approach to database searching that uses only the spectrum to identify the peptide sequence. This method has undergone significant recent improvements, in part due to the integration of machine learning models into the algorithms.  \nThis thesis explores the strengths and weaknesses of many of the current state-ofthe-art de novo peptide identification algorithms through an extensive evaluation. As understanding the underlying data is key to this analysis, a comprehensive survey of the characteristics of tandem mass spectra is included alongside the performance of the algorithms. An alternative machine learning architecture is then proposed to address the weaknesses found. The proposed novel CNN-GNN peptide ion encoding module was able to identify more peptide ions than the encoding modules used by state-of-the-art de novo peptide identification algorithms in all datasets tested. Finally, the utility of artificial data in the context of de novo peptide identification is explored. Artificial spectra were found to be missing critical noise that was present in real data. However, the quantification and introduction of this noise into to artificial spectra increased their similarity to real spectra, significantly improving their potential for use in the training and testing of models. Based on the results of this thesis we recommend specific research avenues for the design and development of the next generation of de novo peptide identification algorithms. This thesis not only demonstrates the challenges facing de novo peptide identification, but also takes the critical first steps toward overcoming them.  \nAcknowledgements  \nFirstly I would like to thank my two supervisors, Dr Enda Howley and Dr Florence Abram. Your advice, support and guidance throughout the last few years has been invaluable. I would also like to thank you for the fun and enthusiasm you brought to the whole process. I hope that this is only the start of our collaboration together.  \nTo those who have passed through Room 307 during my time there, I would like to thank you all. The journey was made much more enjoyable by your friendship and discussion. I was also fortunate to have the support of those in the FEM Lab in Microbiology. The long debates and discussions we had at meetings were always entertaining.  \nFinally I would like to thank my family. To my parents Geraldine and Michael, I am forever grateful for the love and support you have given me. I could not have not have completed this without you. To Marie and John, thank you for always being there forme and looking out for your little brother.  \nContents  \nAbstract ........................................ i  \nAcknowledgements .................................. ii  \nDeclaration ...................................... xviii  \n1 Introduction 1  \n1.1 Motivation ................................... 1  \n1.2 Research Questions ............................... 3  \n1.3 Hypotheses ................................... 3  \n1.4 Thesis Overview ................................ 3  \n2 Background 5  \n2.1 Proteomics ................................... 5  \n2.1.1 Background ............................... 5  \n2.1.2 Methodology ............................","cbCaie4bFBXRTzOy","https://ap.wps.com/l/cbCaie4bFBXRTzOy","pdf",46707171,1,184,"English","en",105,"# Abstract\n# Acknowledgements\n# Declaration\n# Introduction\n## Motivation\n## Research Questions\n## Hypotheses\n## Thesis Overview\n# Background\n## Proteomics\n## Machine learning\n## ML for De Novo\n## Artificial MS/MS Spectra\n# The Impact of Noise and Missing Fragmentation Cleavages on De Novo Peptide Identification Algorithms","[{\"question\":\"Why is database searching insufficient for peptide identification?\",\"answer\":\"On average, about 75% of spectra recovered from experiments remain unidentified when using database searching. This motivates de novo approaches that infer sequences directly from spectra.\"},{\"question\":\"What is the main focus of this thesis?\",\"answer\":\"The thesis evaluates current state-of-the-art de novo peptide identification algorithms, surveys tandem mass spectral characteristics, and proposes a machine learning architecture to address identified weaknesses.\"},{\"question\":\"How does the proposed CNN-GNN module perform compared with existing approaches?\",\"answer\":\"The CNN-GNN peptide ion encoding module identified more peptide ions than encoding modules used by state-of-the-art de novo algorithms across all datasets tested.\"}]","Machine Learning for De Novo Peptide Identification - Doctoral Thesis | PDF",1785727347,464,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"machine-learning-for-de-novo-peptide-identification-doctoral-thesis","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/machine-learning-for-de-novo-peptide-identification-doctoral-thesis/119972/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is database searching insufficient for peptide identification?","Question",{"text":75,"@type":76},"On average, about 75% of spectra recovered from experiments remain unidentified when using database searching. This motivates de novo approaches that infer sequences directly from spectra.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the main focus of this thesis?",{"text":80,"@type":76},"The thesis evaluates current state-of-the-art de novo peptide identification algorithms, surveys tandem mass spectral characteristics, and proposes a machine learning architecture to address identified weaknesses.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the proposed CNN-GNN module perform compared with existing approaches?",{"text":84,"@type":76},"The CNN-GNN peptide ion encoding module identified more peptide ions than encoding modules used by state-of-the-art de novo algorithms across all datasets tested.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]