[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123676-en":3,"doc-seo-123676-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123676,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","Prediction of Melting Temperature of Organic Molecules using Machine Learning - Master’s Thesis - 2023","Accurate prediction of melting points of organic molecules is essential for understanding chemical properties that drive drug behavior and support efficient pharmaceutical discovery. The study addresses the complexity of melting-point prediction by accounting for relationships between enthalpy and entropy shaped by molecular factors such as shape, electronegativity, flexibility, rotatability, and intermolecular bonding. A combined dataset is curated from the Open Notebook Science Dataset and the Cambridge Structure Database and organized into four subsets to analyze feature relevance by bond-forming capacity. Feature engineering is performed using numerical descriptors and embedding features, followed by machine learning training with evaluation via R2 and RMSE, benchmarked against embedding-based models. Analyzing results identifies physical shape descriptors and specific substructural groups as strongly correlated with prediction quality, with additional principal component analysis to explore feature relationships. Results support improved drug screening, formulation, and manufacturing optimization.","Master’s Thesis 2023 30 ECTS  \nFaculty of Science and Technology Professor Kristian Berland  \nPrediction of Melting Temperature of Organic Molecules using Machine Learning  \nAditya Dey  \nMSc Data Science  \nThe page is intentionally left blank.  \nAcknowledgments  \nI would like to express my deep appreciation to my supervisor, Prof. Kristian Berland, for his exceptional guidance, encouragement, and support throughout my research. His invaluable insights, constructive feedback, and unwavering commitment have been instrumental in shaping this project.  \nI would also like to extend my sincere gratitude to Seyedmojtaba Seyedraouﬁ and Elin Sødahl for their invaluable assistance, generous support, and insightful guidance. Their contributions have been immensely helpful in advancing my research and achieving my goals.  \nFinally, I would like to thank my family and friends for their unconditional love, unwavering support, and understanding throughout this journey. Their constant encouragement and belief in me have kept me motivated and inspired me to pursue my passion.  \nAditya Dey  \nÅs, May 15th 2023  \nThe page is intentionally left blank.  \nAbstract  \nAccurate prediction of the melting point of oral drugs is crucial for understanding their chemical properties. Early identiﬁcation of these properties aids in the screening of potential drugs, thereby saving resources in the pharmaceutical industry’s discovery and manufacturing processes. The prediction of organic molecules is a complex task due to many factors that aﬀect entropy and enthalpy forces within a molecule, which are dependent on various factors like shape, electronegativity, ﬂexibility, rotatability, intermolecular bonding, etc.  \nIn this study, we curated a combined dataset of organic molecules, extracted from the Open Notebook Science Dataset and Cambridge Structure Database. The dataset consists of molecules composed of carbon, oxygen, nitrogen, sulfur, phosphorous, and halogens, exhibiting a wide range of melting point temperatures and molecules with complex structures. To gain insights into the signiﬁcance of each feature and its contribution to melting point prediction, we divided the combined dataset into four subsets based on the number of bonds an atom can form.  \nWe perform feature engineering on these datasets by studying the physical and chemical properties known to impact melting points. Numerical features were derived from the molecules, capturing relevant information. Additionally, we utilized embedding features without any modiﬁcations.  \nMachine learning models were trained using both numerical and embedding features, with the accuracy evaluated through R2 scores and root mean squared error values. We set the model trained on embedding features as a benchmark for our model and features to surpass. Our machine learning models exhibited good performance, outperforming the benchmark and achieving good prediction accuracy.  \nFurthermore, we conducted an in-depth analysis of the results to assess the impact of individual features on the models. We observed physical shape features and the presence of speciﬁc substructural groups exhibited a strong correlation with melting point prediction. To explore the relationship between features, we performed a principal component analysis.  \nThe ﬁndings of this study have important implications for drug development, formulation, and optimization of manufacturing processes. Accurate prediction of melting points enhances drug screening procedures and aids in the design of eﬀective pharmaceutical products.  \nThe codes are available in github.  \nThe page is intentionally left blank.  \nContents  \n1 Introduction 16  \n1. 1 Aims of this Master’s Thesis . . . . . . . . . . . . . . . . . . . . . . . . . . 20  \n2 Datasets 21  \n2.1 Cambridge Structure Database ........................ 21  \n2.2 Open Notebook Science Melting Point Dataset ................ 22  \n2.3 Combined Dataset . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .","cbCaivoHxwjbyYvB","https://ap.wps.com/l/cbCaivoHxwjbyYvB","pdf",6355682,1,99,"English","en",105,"# Introduction\n## Aims of this Master’s Thesis\n# Datasets\n## Cambridge Structure Database\n## Open Notebook Science Melting Point Dataset\n## Combined Dataset\n## Segregation of Combined Dataset\n# Feature Engineering\n# Machine Learning Theory\n## Support Vector Machines\n## Random Forests\n## AdaBoost\n## XGBoost\n## Artificial Neural Network\n# Machine Learning Training\n## Feature Selection\n## Hyperparameter Selection\n## Training, Validation and Testing\n## Performance Metrics\n# Results\n# Discussion\n## Model Analysis\n## Future Improvements\n# Conclusion\n# Appendices\n## Scatter Plots\n## Partial Dependence Plot","[{\"question\":\"Why is predicting melting temperature of organic (oral) drugs important?\",\"answer\":\"Accurate melting-point prediction helps reveal chemical properties that matter for drug behavior. Early identification supports screening of candidate drugs and reduces resources required in discovery and manufacturing.\"},{\"question\":\"How was the dataset built and organized for this study?\",\"answer\":\"The dataset combines molecules from the Open Notebook Science Dataset and the Cambridge Structure Database. It is split into four subsets based on the number of bonds an atom can form to enable feature-significance analysis.\"},{\"question\":\"What types of features and models were used, and how were they evaluated?\",\"answer\":\"The work uses numerical features derived from molecular properties and embedding features without modifications. Machine learning models are trained using both feature types and evaluated with R2 scores and root mean squared error (RMSE), with embedding-based models as a benchmark.\"}]","Prediction of Melting Temperature of Organic Molecules using Machine Learning - Master’s Thesis - 2023 | PDF",1785817966,249,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"prediction-of-melting-temperature-of-organic-molecules-using-machine-learning-masters-thesis-2023","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/prediction-of-melting-temperature-of-organic-molecules-using-machine-learning-masters-thesis-2023/123676/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is predicting melting temperature of organic (oral) drugs important?","Question",{"text":75,"@type":76},"Accurate melting-point prediction helps reveal chemical properties that matter for drug behavior. Early identification supports screening of candidate drugs and reduces resources required in discovery and manufacturing.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How was the dataset built and organized for this study?",{"text":80,"@type":76},"The dataset combines molecules from the Open Notebook Science Dataset and the Cambridge Structure Database. It is split into four subsets based on the number of bonds an atom can form to enable feature-significance analysis.",{"name":82,"@type":73,"acceptedAnswer":83},"What types of features and models were used, and how were they evaluated?",{"text":84,"@type":76},"The work uses numerical features derived from molecular properties and embedding features without modifications. Machine learning models are trained using both feature types and evaluated with R2 scores and root mean squared error (RMSE), with embedding-based models as a benchmark.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]