[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-125642-en":3,"doc-seo-125642-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},125642,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","Machine Learning Small Molecule Properties in Drug Discovery - Comprehensive Overview","Machine learning offers a promising pathway for predicting small-molecule properties in drug discovery. The review consolidates recent advances in methods for diverse targets, covering binding affinities, solubility, and ADMET metrics such as absorption, distribution, metabolism, excretion, and toxicity. It surveys widely used datasets and molecular representations including chemical fingerprints and graph-based neural networks. It also addresses challenges in multi-property prediction and optimization across hit-to-lead and lead stages, and evaluates interpretability approaches for supporting critical decision-making. The survey emphasizes that performance depends strongly on training data quality and highlights the need for standardized benchmarks and comparable metrics.","arXiv :2308 . 12354v1 [ q-bio .BM] 2 Aug 2023  \nMachine Learning Small Molecule Properties in Drug Discovery  \nNikolai Schapin∗a,b, Maciej Majewskia , Alejandro Varelaa , Carlos Arroniza , and  \nGianni De Fabritiis†b,c,d  \naAcellera Labs, C/ Doctor Trueta 183, 08005 Barcelona, Spain b Computational Science Laboratory, Universitat Pompeu Fabra, PRBB, C/ Doctor  \nAiguader 88, 08003 Barcelona, Spain  \nc Instituci´o Catalana de Recerca i Estudis Avan¸cats (ICREA), Passeig Llu´ıs Companys  \n23, 08010 Barcelona, Spain  \ndAcellera, Devonshire House 582 Honeypot Lane Stanmore Middlesex,HA7 1JS  \nUnited Kingdom  \nAugust 25, 2023  \nAbstract  \nMachine learning (ML) is a promising approach for predicting small molecule properties in drug discovery. Here, we provide a comprehensive overview of various ML methods introduced for this purpose in recent years. We review a wide range of properties, including binding affinities, solubility, and ADMET (Absorption, Distribution, Metabolism, Excretion, and Toxicity) . We discuss existing popular datasets and molecular descriptors and embeddings, such as chemical fingerprints and graph-based neural networks. We highlight also challenges of predicting and optimizing multiple properties during hit-to-lead and lead optimization stages of drug discovery and explore briefly possible multi-objective optimization techniques that can be used to balance diverse properties while optimizing lead candidates. Finally, techniques to provide an understanding of model predictions, especially for critical decision-making in drug discovery are assessed. Overall, this review provides insights into the landscape of ML models for small molecule property predictions in drug discovery. So far, there are multiple diverse approaches, but their performances are often comparable. Neural networks, while more flexible, do not always outperform simpler models. This shows that the availability of high-quality training data remains crucial for training accurate models and there is a need for standardized benchmarks, additional performance metrics, and best practices to enable richer comparisons between the different techniques and models that can sheda better light on the differences between the many techniques.  \nKeywords—molecular property prediction, ADMET prediction models, binding affinity prediction models, physicochemical properties prediction models, computational methods in drug discovery  \n1 Introduction  \nEarly stage, preclinical drug discovery is a step-wise process, where at each stage, hit molecules are required to meet certain criteria to ensure their efficacy and quality before proceeding to the next stage. This results in a series of molecular properties that need to be optimized. In order to do this, they need to be measured, which traditionally is done through wet-lab experiments that are costly and time-consuming. The estimated R&D expenditure is around 41 billion euros in Europe and 83 billion US dollars in the USA with R&D costs per drug ranging around 1-2 billion US dollars [1, 2] . This process takes on average 10-13 years [1, 2], and only 1 out of 10 000 substances tested [2] will pass through all the stages to become a new successfully marketed drug. While most of the cost and two-thirds of the time are linked to the stages of clinical trials, most of the candidate molecules fail during these stages [3] . The main reasons [4] for failure are low efficacy of the drug, high toxicity,  \n∗ corresponding author, e-mail address: [n.shapin@acellera.com](n.shapin@acellera.com)  \n†corresponding author, e-mail address: [g.defabritiis@acellera.com](g.defabritiis@acellera.com)  \nor commercial reasons. The first two are often the result of unsuccessful or insufficient establishment of key molecular properties of the hit and lead candidates during the early stages of drug discovery.  \nVarious resource-efficient computational techniques have been developed which all fall under the group of computer-aided drug design met","cbCaihyUg6GAdHMN","https://ap.wps.com/l/cbCaihyUg6GAdHMN","pdf",1026566,1,46,"English","en",105,"# Introduction\n## Early-stage drug discovery and bottlenecks\n## Computational approaches for property estimation\n## Docking, simulations, and empirical scoring functions","[{\"question\":\"Why is machine learning valuable for small-molecule property prediction in drug discovery?\",\"answer\":\"Machine learning enables efficient estimation of molecular properties for large compound sets without relying solely on costly and time-consuming wet-lab experiments.\"},{\"question\":\"Which types of properties are covered in the review?\",\"answer\":\"The review addresses binding affinities, solubility, and ADMET-related properties including absorption, distribution, metabolism, excretion, and toxicity.\"},{\"question\":\"What are the key challenges when optimizing multiple properties during lead discovery?\",\"answer\":\"Predicting and optimizing multiple properties across hit-to-lead and lead optimization stages is difficult, motivating multi-objective approaches to balance competing property requirements.\"}]","Machine Learning Small Molecule Properties in Drug Discovery - Comprehensive Overview | PDF",1785900376,116,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"machine-learning-small-molecule-properties-in-drug-discovery-comprehensive-overview","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/machine-learning-small-molecule-properties-in-drug-discovery-comprehensive-overview/125642/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is machine learning valuable for small-molecule property prediction in drug discovery?","Question",{"text":75,"@type":76},"Machine learning enables efficient estimation of molecular properties for large compound sets without relying solely on costly and time-consuming wet-lab experiments.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which types of properties are covered in the review?",{"text":80,"@type":76},"The review addresses binding affinities, solubility, and ADMET-related properties including absorption, distribution, metabolism, excretion, and toxicity.",{"name":82,"@type":73,"acceptedAnswer":83},"What are the key challenges when optimizing multiple properties during lead discovery?",{"text":84,"@type":76},"Predicting and optimizing multiple properties across hit-to-lead and lead optimization stages is difficult, motivating multi-objective approaches to balance competing property requirements.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]