[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120120-en":3,"doc-seo-120120-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120120,8796095461564,"Liam","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Machine learning for molecular property prediction and drug safety - A broad perspective on using deep learning to predict acid dissociation constants - Master’s thesis","Utilizing machine learning methods for the prediction of acid dissociation (pKa) values of compounds holds great significance, as pKa is an important parameter optimized frequently in drug discovery. Accurate pKa prediction can provide insights into other molecular properties and support compound design. Using internal AstraZeneca data, classical ML with molecular descriptors and deep learning models were developed and compared. Graph neural networks outperformed tree-based methods, showed reasonable performance for acidic and basic pKa, and generalized well to novel compounds. Evaluation on public datasets yielded lower accuracies due to diverse sources and high experimental variability.","Machine learning for molecular property prediction and drug safety  \nA broad perspective on using deep learning to predict acid dissociation constants  \nMaster’s thesis in Applied Data Science  \nKINGA JENEI  \nDepartment of Computer Science and Engineering CHALMERS UNIVERSITY OF TECHNOLOGY UNIVERSITY OF GOTHENBURG  \nGothenburg, Sweden 2023  \nMaster’s thesis 2023  \nMachine learning for molecular property prediction and drug safety  \nA broad perspective on using deep learning to predict acid  \ndissociation constants  \nKINGA JENEI  \nDepartment of Computer Science and Engineering Division of Data Science and AI Chalmers University of Technology University of Gothenburg Gothenburg, Sweden 2023  \nMachine learning for molecular property prediction and drug safety  \nA broad perspective on using deep learning to predict acid dissociation constants KINGA JENEI  \n© KINGA JENEI, 2023 .  \nSupervisor: Rocío Mercado, Department of Computer Science and Engineering  \nCompany supervisor: Vigneshwari Subramanian, AstraZeneca  \nExaminer: Ola Engkvist, Department of Computer Science and Engineering  \nMaster’s Thesis 2023  \nDepartment of Computer Science and Engineering Division of Data Science and AI  \nChalmers University of Technology and University of Gothenburg SE-412 96 Gothenburg  \nTelephone +46 31 772 1000  \nTypeset in LATEX  \nGothenburg, Sweden 2023  \nMachine learning for molecular property prediction and drug safety  \nA broad perspective on using deep learning to predict acid dissociation constants KINGA JENEI  \nDepartment of Computer Science and Engineering  \nChalmers University of Technology and University of Gothenburg  \nAbstract  \nUtilizing machine learning methods for the prediction of acid dissociation (pKa ) values of compounds holds great signiﬁcance, as pKa is an important parameter, optimized frequently in drug discovery. Accurate prediction of pKa values could potentially provide valuable insights on other molecular properties and thereby support compound design. In an attempt to extend the scope of pKa prediction, we have created several machine learning models utilizing internal AstraZeneca data. We explored both classical ML approaches with diﬀerent molecular descriptors, and deep learning methods. The results showed that graph neural network based models outperform tree based methods and yielded reasonable predictions for both acidic and basic pKa values. Through the implementation of several data splitting strategies, we have substantiated that the models hold the potential to generalize well to novel compounds and outperform state of the art methods. Besides evaluating the models on diﬀerent splits of the internal data, their performance was also assessed on public datasets. This yielded comparatively lower accuracies which can be attributed to the collation of data from diverse sources and the high experimental variability of the publicly available data.  \nKeywords: Molecular property prediction, Acid dissociation constant, pKa, Machine learning, Graph Neural Networks, Molecular descriptors, Drug Discovery.  \nAcknowledgements  \nI would like to express my deepest gratitude to my company supervisor, Vigneshwari Subramanian, for her guidance, invaluable feedback and encouragement throughout the project. Additionally, this endeavor would not have been possible without my university supervisor, Rocío Mercado, whose knowledge, ideas and advice helped shape this project. I would also like to express my appreciation to Emma Evertsson and Susanne Winiwarter for their support and assistance, and whose expertise has been a great contribution to this work. I have valued the time spent at AstraZeneca and would like to thank everyone who I have met and had the opportunity to work with throughout the last six months. I am also grateful for the input of my examiner, Ola Engkvist, and for the feedback and editing help of my opponents. Last but not least, I would like to thank my family and friends for supporting, motivating and believing i","cbCaiiLXB3iGAf2D","https://ap.wps.com/l/cbCaiiLXB3iGAf2D","pdf",1587984,1,62,"English","en",105,"# Introduction\n## Background\n## Goals and challenges\n## Thesis outline\n# Theory\n## Representing Molecules\n## Models\n## Model training\n## Evaluation metrics\n# Methods\n## Workflow\n## Data\n## Model training and evaluation\n# Results\n## Single task models\n## Multitask models\n## Comparison to state of the art\n## Potential to predict new modalities\n# Conclusion\n## Discussion","[{\"question\":\"Why is accurate pKa prediction important in drug discovery?\",\"answer\":\"pKa is an important, frequently optimized parameter in drug discovery. Accurate predictions can also provide insights into other molecular properties that support compound design.\"},{\"question\":\"What modeling approaches were explored for pKa prediction?\",\"answer\":\"The thesis compares classical machine learning models using different molecular descriptors with deep learning methods.\"},{\"question\":\"How did graph neural network models perform compared to tree-based models?\",\"answer\":\"Graph neural network based models outperformed tree based methods and produced reasonable predictions for both acidic and basic pKa values.\"}]","Machine learning for molecular property prediction and drug safety - A broad perspective on using deep learning to predict acid dissociation constants - Master’s thesis | PDF",1785728310,156,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"machine-learning-for-molecular-property-prediction-and-drug-safety-a-broad-perspective-on-using-deep-learning-to-predict-acid-dissociation-constants-masters-thesis","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/machine-learning-for-molecular-property-prediction-and-drug-safety-a-broad-perspective-on-using-deep-learning-to-predict-acid-dissociation-constants-masters-thesis/120120/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is accurate pKa prediction important in drug discovery?","Question",{"text":75,"@type":76},"pKa is an important, frequently optimized parameter in drug discovery. Accurate predictions can also provide insights into other molecular properties that support compound design.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What modeling approaches were explored for pKa prediction?",{"text":80,"@type":76},"The thesis compares classical machine learning models using different molecular descriptors with deep learning methods.",{"name":82,"@type":73,"acceptedAnswer":83},"How did graph neural network models perform compared to tree-based models?",{"text":84,"@type":76},"Graph neural network based models outperformed tree based methods and produced reasonable predictions for both acidic and basic pKa values.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]