[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-125213-en":3,"doc-seo-125213-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},125213,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","Machine Learning for Toxicity Prediction Using Chemical Structures: Pillars for Success in the Real World","Machine learning (ML) increasingly supports the prediction of molecular properties and toxicity in drug discovery, but experimental toxicity endpoints remain difficult to translate in vivo because human and animal studies require substantial resources, limiting available data. ML may complement or replace parts of experimental workflows depending on project goals. Yet ML adoption carries risks from biased training data, mismatched model choices, and weak building or validation, leading to inaccurate predictions and poor decision-making. This perspective argues for improving predictive validity by focusing on defined datasets from small-molecule structures and five success pillars: dataset selection, structural representations, model algorithm, model validation, and translation into decisions.","This article is licensed under CC-BY 4.0   \n[pubs.acs.org/crt](pubs.acs.org/crt)  Review   \nMachine Learning for Toxicity Prediction Using Chemical Structures: Pillars for Success in the Real World  \nSrijit Seal, * Manas Mahale, Miguel García-Ortegón, Chaitanya K. Joshi, Layla Hosseini-Gerami, Alex Beatson, Matthew Greenig, Mrinal Shekhar, Arijit Patra, Caroline Weis, Arash Mehrjou, Adrien Badré, Brianna Paisley, Rhiannon Lowe, Shantanu Singh, Falgun Shah, Bjarki Johannesson, Dominic Williams, David Rouquie, Djork-Arné Clevert, Patrick Schwab, Nicola Richmond, Christos A. Nicolaou, Raymond J. Gonzalez, Russell Naven, Carolin Schramm, Lewis R Vidler, Kamel Mansouri, W. Patrick Walters, Deidre Dalmas Wilk, Ola Spjuth, * Anne E. Carpenter, * and Andreas Bender *  \n Cite This: Chem. Res. Toxicol. 2025, 38, 759−807  \nRead Online  \n\n|  |  |  |  |\n| --- | --- | --- | --- |\n| ACCESS   | Metrics & More |  |  Article Recommendations |\n\nABSTRACT: Machine learning (ML) is increasingly valuable for predicting molecular properties and toxicity in drug discovery. However, toxicity-related end points have always been challenging to evaluate experimentally with respect to in vivo translation due to the required resources for human and animal studies; this has impacted data availability in the field. ML can augment or even potentially replace traditional experimental processes depending on the project phase and specific goals of the prediction. For instance, models can be used to select promising compounds for on-target effects or to deselect those with undesirable characteristics (e.g., off-target or ineffective due to unfavorable pharmacokinetics). However, reliance on ML is not without risks, due to biases stemming from nonrepresentative training data, incompatible choice of algorithm to represent the underlying data, or poor model building and validation approaches. This might lead to inaccurate predictions, misinterpretation of the confidence in ML predictions, and ultimately suboptimal decision-making. Hence, understanding the predictive validity of ML models is of utmost importance to enable faster drug development timelines while improving the quality of decisions. This perspective emphasizes the need to enhance the understanding and application of machine learning models in drug discovery, focusing on well-defined data sets for toxicity prediction based on small molecule structures. We focus on five crucial pillars for success with ML-driven molecular property and toxicity prediction: (1) data set selection, (2) structural representations, (3) model algorithm, (4) model validation, and (5) translation of predictions to decision-making. Understanding these key pillars will foster collaboration and coordination between ML researchers and toxicologists, which will help to advance drug discovery and development.  \n■ INTRODUCTION  \nIn recent years, machine learning (ML) approaches for toxicity prediction using chemical structures and sometimes additional data sources have attracted widespread interest, particularly in drug discovery.1 There are constant innovations in ML for investigating biological systems and understanding their interactions with drugs, resulting in therapeutic activity and/or  \n\n| Received: January 26, 2025\u003Cbr>Revised: March 24, 2025\u003Cbr>Accepted: March 25, 2025\u003Cbr>Published: May 2, 2025 |  |\n| --- | --- |\n\n© 2025 The Authors. Published by American Chemical Society  \n759  \n[https://doi.org/10.1021/acs.chemrestox.5c00033](https://doi.org/10.1021/acs.chemrestox.5c00033)  \nChem. Res. Toxicol. 2025, 38, 759−807  \nadverse outcomes. Still, these improvements must be reflected in practice. Currently, the costs of bringing a drug to market are increasing,2 while the overall success rates in clinical drug development remain poor.3 ML models can be trained on data from empirical assays to predict the properties of compounds from molecular structure (see Box 1 for standard definitions for ML related to predictive models for molecula","cbCaicMiXLwhRijO","https://ap.wps.com/l/cbCaicMiXLwhRijO","pdf",11086170,1,49,"English","en",105,"# Introduction\n## QSAR/QSPR and toxicity prediction workflow\n## Predicting efficacy, toxicity, and ADME/PK properties","[{\"question\":\"Why is ML toxicity prediction challenging to rely on for in vivo translation?\",\"answer\":\"Toxicity endpoints are difficult to evaluate experimentally with respect to in vivo translation, because human and animal studies demand substantial resources, which constrains data availability.\"},{\"question\":\"What are the key risks when using ML for toxicity prediction?\",\"answer\":\"Risks include biases from nonrepresentative training data, choosing algorithms that do not fit how the data is represented, and poor model building and validation—leading to inaccurate predictions and misjudging confidence.\"},{\"question\":\"What are the five pillars for success in ML-driven toxicity prediction?\",\"answer\":\"The perspective highlights five pillars: (1) dataset selection, (2) structural representations, (3) model algorithm, (4) model validation, and (5) translation of predictions into decision-making.\"}]","Machine Learning for Toxicity Prediction Using Chemical Structures: Pillars for Success in the Real World | PDF",1785897489,123,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"machine-learning-for-toxicity-prediction-using-chemical-structures-pillars-for-success-in-the-real-world","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/machine-learning-for-toxicity-prediction-using-chemical-structures-pillars-for-success-in-the-real-world/125213/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is ML toxicity prediction challenging to rely on for in vivo translation?","Question",{"text":75,"@type":76},"Toxicity endpoints are difficult to evaluate experimentally with respect to in vivo translation, because human and animal studies demand substantial resources, which constrains data availability.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What are the key risks when using ML for toxicity prediction?",{"text":80,"@type":76},"Risks include biases from nonrepresentative training data, choosing algorithms that do not fit how the data is represented, and poor model building and validation—leading to inaccurate predictions and misjudging confidence.",{"name":82,"@type":73,"acceptedAnswer":83},"What are the five pillars for success in ML-driven toxicity prediction?",{"text":84,"@type":76},"The perspective highlights five pillars: (1) dataset selection, (2) structural representations, (3) model algorithm, (4) model validation, and (5) translation of predictions into decision-making.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]