[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122753-en":3,"doc-seo-122753-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122753,8796095360427,"Lucas Martin","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","Accelerating the Design-Make-Test cycle of Drug Discovery with Machine Learning","Drug discovery follows a design-make-test loop in which proposed compounds are synthesised and tested for bioactivity, and each round informs the next. The long preclinical timeline stems from difficulties at every step in the pipeline. This thesis investigates how machine learning can accelerate the design-make-test cycle, enabling faster hit finding, more effective hit-to-lead optimisation, and improved reaction and bioactivity modelling. It develops unsupervised methods, learning-to-rank frameworks, interpretable reaction workflows, and assays using crude mixtures.","Accelerating the Design-Make-Test cycle of Drug Discovery with Machine  \nLearning  \nWilliam McCorkindale  \nCavendish Laboratory, Department of Physics University of Cambridge  \nSupervisor: Dr. Alpha Lee  \nSt. John’s College October 2023  \nDeclaration  \nI hereby declare that except where speciﬁc reference is made to the work of others, the contents of this dissertation are original and have not been submitted in whole or in part for consideration for any other degree or qualiﬁcation in this, or any other university. This dissertation is my own work and contains nothing which is the outcome of work done in collaboration with others, except as speciﬁed in the Preface and Acknowledgements. This dissertation contains fewer than 65,000 words including appendices, bibliography, footnotes, tables and equations and has fewer than 150 ﬁgures.  \nThe research described in this thesis was performed between October 2019 and December 2022, and was supervised by Dr Alpha A. Lee.  \nWilliam McCorkindale October 2023  \nAccelerating the Design-Make-Test cycle of Drug Discovery  \nwith Machine Learning  \nWilliam McCorkindale  \nAbstract  \nDrug discovery follows a design-make-test cycle of proposing drug compounds, synthesising them, and measuring their bioactivity, which informs the next cycle of compound designs. The challenges associated with each step lead to the long timeline of preclinical pharmaceutical development. This thesis focuses on how we can use machine learning tools to accelerate the design-make-test cycle for faster drug discovery.  \nWe begin with the design of new compounds, looking at the initial stage of fragment-based hit ﬁnding where only the 3D coordinates of fragment-protein complexes are available. The standard approach is to “grow” or “merge” nearby fragments based on their binding modes, but fragments typically have low afﬁnity so the road to potency is often long and fraught with false starts. Instead, we can reframe fragment-based hit discovery as a denoising problemidentifying signiﬁcant pharmacophore distributions from an “ensemble” of fragments amid noise due to weak binders-and employ an unsupervised machine learning method to tackle this problem. We construct a model that screens potential molecules by evaluating whether they recapitulate those fragment-derived pharmacophore distributions. We show that this approach outperforms docking in distinguishing active compounds from inactive ones on historical data. Further, we prospectively ﬁnd novel hits for SARS-CoV-2 Mpro and the Mac1 domain of SARS-CoV-2 non-structural protein 3 by screening a library of 1 billion molecules.  \nAfter identifying hit compounds, we enter the hit-to-lead stage where we wish to optimise their molecular structures to improve bioactivity. Framing bioactivity modelling as active/inactive classiﬁcation would not allow us to rank compounds based on predicted bioactivity improvement, while the low number of active compounds and the measurement noise make a regression approach challenging. We overcome this challenge with a learning-to-rank framework via a classiﬁer that predicts whether a compound is more or less active than another using the difference in molecular descriptors between the molecules as input. This allows us to make use of inactive data, and threshold the bioactivity differences above measurement noise. Validation on retrospective data for Mpro shows that we can outperform docking on ranking ligands, and we prospectively screen a library of 8.8M molecules and arrive at a potent compound with a novel scaffold.  \nThroughout the entire course of drug discovery, one needs to ﬁnd a synthesis route to actually make the molecule. An exciting approach is to use deep learning models trained on  \npatent reaction databases, but they suffer from being opaque black boxes. It is neither clear if the models are making correct predictions because they inferred the salient chemistry, nor is it clear which training data they are relying on to reach ","cbCaiqNzU1Bgo6Vs","https://ap.wps.com/l/cbCaiqNzU1Bgo6Vs","pdf",24899550,1,138,"English","en",105,"# Abstract\n## Design: fragment-based hit finding with unsupervised ML\n## Hit-to-lead: learning-to-rank bioactivity modelling\n## Synthesis route: interpretable deep learning for reaction prediction\n## Testing: bioactivity from crude reaction mixtures","[{\"question\":\"How does the thesis aim to accelerate the drug discovery design-make-test cycle?\",\"answer\":\"By applying machine learning to multiple stages, including designing compounds, improving hit-to-lead optimisation, interpreting reaction prediction models for synthesis, and increasing throughput for bioactivity measurement.\"},{\"question\":\"What method is used for fragment-based hit finding when only limited structural data is available?\",\"answer\":\"The work reframes fragment-based discovery as a denoising problem and uses an unsupervised machine learning approach to identify significant pharmacophore distributions from an ensemble of fragments amid noise.\"},{\"question\":\"How does the thesis address the interpretability issues of deep learning reaction models?\",\"answer\":\"It develops a workflow for quantitatively interpreting state-of-the-art deep learning reaction prediction models, including analysis of chemically selective reactions to reveal correct reasoning, counterintuitive predictions, and dataset-bias-driven “Clever Hans” behavior.\"}]","Accelerating the Design-Make-Test cycle of Drug Discovery with Machine Learning | PDF",1785812715,348,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"accelerating-the-design-make-test-cycle-of-drug-discovery-with-machine-learning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/accelerating-the-design-make-test-cycle-of-drug-discovery-with-machine-learning/122753/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How does the thesis aim to accelerate the drug discovery design-make-test cycle?","Question",{"text":75,"@type":76},"By applying machine learning to multiple stages, including designing compounds, improving hit-to-lead optimisation, interpreting reaction prediction models for synthesis, and increasing throughput for bioactivity measurement.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What method is used for fragment-based hit finding when only limited structural data is available?",{"text":80,"@type":76},"The work reframes fragment-based discovery as a denoising problem and uses an unsupervised machine learning approach to identify significant pharmacophore distributions from an ensemble of fragments amid noise.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the thesis address the interpretability issues of deep learning reaction models?",{"text":84,"@type":76},"It develops a workflow for quantitatively interpreting state-of-the-art deep learning reaction prediction models, including analysis of chemically selective reactions to reveal correct reasoning, counterintuitive predictions, and dataset-bias-driven “Clever Hans” behavior.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]