[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118278-en":3,"doc-seo-118278-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118278,7971461740886,"Theodore","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","Probabilistic machine learning algorithms for molecule discovery - Doctoral dissertation","Discovering new molecules enables progress in health, agriculture, energy, and more, but experimental testing is limited compared with the enormous space of possible molecular structures. The thesis frames molecule discovery as a Bayesian optimisation problem that selects candidate molecules using both existing knowledge and expected information gain from each test. Because molecules are discrete, practical Bayesian optimisation requires new probabilistic machine learning algorithms. The work proposes multiple methods using Gaussian processes with deep neural network kernels, Tanimoto random features for large datasets, and a retro-fallback model for retrosynthesis feasibility.","Probabilistic machine learning algorithms for molecule discovery  \nAustin James Tripp  \nDepartment of Engineering  \nUniversity of Cambridge  \nThis dissertation is submitted for the degree of Doctor of Philosophy  \nGirton College April 2024  \nDeclaration  \nThis thesis is the result of my own work and includes nothing which is the outcome of work done in collaboration except as declared in the preface and specified in the text. It is not substantially the same as any work that has already been submitted, or is being concurrently submitted, for any degree, diploma or other qualification at the University of Cambridge or anyother University or similar institution except as declared in the preface and specified in the text. It does not exceed the prescribed word limit for the relevant Degree Committee.  \nAustin James Tripp April 2024  \nAbstract  \nDiscovering new molecules empowers humanity to solve problems in health, agriculture, energy, and more. The key challenge of molecule discovery is that the space of all possible molecules is vastly larger than the amount of molecules which we can test experimentally with our limited resources. Given this challenge, arguably the best option is to judiciously select which molecules to test based on both our current knowledge and our expectation of what information will be gained from each test.  \nIn machine learning, this approach is typically called Bayesian optimisation and has been studied for many other problems, such as tuning hyperparameters of machine learning models. Although in principle Bayesian optimisation can be straightforwardly applied to the problem of discovering new molecules, the discrete nature of molecules means that new models and algorithms are needed to make Bayesian optimisation work in practice.  \nThis thesis presents various probabilistic machine learning algorithms which could be used within a Bayesian optimisation loop to discover new molecules. Latent space optimisation with weighted retraining (chapter 3) and adaptive deep kernel fitting with implicit function theorem (chapter 4) are both algorithms which use a Gaussian process with a deep neural network kernel function to model the relationship between molecular structure and some property of interest. Tanimoto random features (chapter 5) allows an established cheminformatics model tobe applied (approximately) to large datasets. Finally, retro-fallback (chapter 6) uses a novel probabilistic formulation of the retrosynthesis problem to estimate whether a molecule can be synthesized, and thereby determine whether it should be considered by Bayesian optimisation. Together, these algorithms form a suite of tools which could be used to discover new molecules automatically and intelligently.  \nAcknowledgements  \nSo many people made my PhD journey possible. First and foremost I would like to thank my PhD supervisor, Miguel. I think Miguel strikes a very good balance in his supervision style. He gave me the freedom to explore my own ideas right from the start of my PhD, and would always be encouraging, while also being upfront about potential flaws in my ideas. This made me feel responsible for my own research, yet also supported. I think this was a very good way to learn, and I am very grateful for this. I’m also very grateful to my advisor, Adrian Weller, for his support and guidance when I needed it.  \nI worked with many great collaborators during my PhD. Erik Daxberger: I appreciated your perpetual calm. Wenlin Chen: I appreciated your hard work and critical thinking. Gregor Simm and Marwin Segler: I appreciated your mentorship and impressive chemistry knowledge. Sergio Bacallado: I appreciated your mentorship, enthusiasm, patience, and ability to explain things clearly. Miguel García-Ortegón: I appreciated your insights into the biology behind drug discovery. Krzysztof Maziarz: I appreciated your mentorship, impressive coding skills, and legendary attention to detail when reviewing pull requests during my interns","cbCaijqnA6enqGMf","https://ap.wps.com/l/cbCaijqnA6enqGMf","pdf",3088549,1,182,"English","en",105,"# Introduction\n## Definition and challenges of “molecule discovery”\n## Why consider probabilistic machine learning?\n## Contributions of thesis\n## Guide to reading this thesis\n# Background: algorithmic tools for molecule discovery","[{\"question\":\"What main problem does the thesis address in molecule discovery?\",\"answer\":\"The thesis addresses the mismatch between the vast space of possible molecules and the limited number that can be tested experimentally.\"},{\"question\":\"Why does Bayesian optimisation need new algorithms for molecule discovery?\",\"answer\":\"Because molecules are discrete, applying Bayesian optimisation requires probabilistic models and algorithms that can work effectively in this discrete domain.\"},{\"question\":\"What algorithmic components does the thesis propose for the Bayesian optimisation loop?\",\"answer\":\"It presents several probabilistic tools, including latent space optimisation with weighted retraining, adaptive deep kernel fitting, Tanimoto random features for large datasets, and a retro-fallback formulation for retrosynthesis feasibility.\"}]","Probabilistic machine learning algorithms for molecule discovery - Doctoral dissertation | PDF",1785682779,459,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"probabilistic-machine-learning-algorithms-for-molecule-discovery-doctoral-dissertation","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/probabilistic-machine-learning-algorithms-for-molecule-discovery-doctoral-dissertation/118278/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What main problem does the thesis address in molecule discovery?","Question",{"text":75,"@type":76},"The thesis addresses the mismatch between the vast space of possible molecules and the limited number that can be tested experimentally.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why does Bayesian optimisation need new algorithms for molecule discovery?",{"text":80,"@type":76},"Because molecules are discrete, applying Bayesian optimisation requires probabilistic models and algorithms that can work effectively in this discrete domain.",{"name":82,"@type":73,"acceptedAnswer":83},"What algorithmic components does the thesis propose for the Bayesian optimisation loop?",{"text":84,"@type":76},"It presents several probabilistic tools, including latent space optimisation with weighted retraining, adaptive deep kernel fitting, Tanimoto random features for large datasets, and a retro-fallback formulation for retrosynthesis feasibility.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]