[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123279-en":3,"doc-seo-123279-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123279,687197207057,"Sage","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Improved Machine Learning Predictions of EC50s Using Uncertainty Estimation from Dose-Response Data","Early-stage drug design often compresses experimental results into a single molecule-level EC50 via curve fitting, losing information about curve-fit reliability. This study adds a fit-quality metric to machine learning QSAR models so each molecule’s EC50 is accompanied by a reliability-aware uncertainty signal. Using 40 PubChem (public) and BASF (private) dose-response datasets, fit-quality incorporation improves predictive performance on 31 datasets. Tested methods include random forests with parametric bootstrap, weighted random forests, variable output smearing random forests, and weighted support vector regression, achieving up to a 22% RMSE reduction.","This article is licensed under CC-BY 4.0   \n[pubs.acs.org/jcim](pubs.acs.org/jcim)  Article   \nImproved Machine Learning Predictions of EC50s Using Uncertainty Estimation from Dose−Response Data  \nPublished as part of Journal of Chemical Information and Modeling special issue “Chemical Compound Space Exploration by Multiscale High-Throughput Screening and Machine Learning”.  \nHugo Bellamy, * Joachim Dickhaut, and Ross D. King  \n Cite This: J. Chem. Inf. Model. 2025, 65, 5623−5634  \nRead Online  \n\n|  |  |  |  |\n| --- | --- | --- | --- |\n| ACCESS   | Metrics & More |  |  Article Recommendations |\n\nABSTRACT: In early-stage drug design, machine learning models often rely on compressed representations of data, where raw experimental results are distilled into a single metric per molecule through curve fitting. This process discards valuable information about the quality of the curve fit. In this study, we incorporated a fit-quality metric into machine learning models to capture the reliability of metrics for individual molecules. Using 40 data sets from PubChem (public) and BASF (private), we demonstrated that including this quality metric can significantly improve predictive performance without additional experiments. Four methods were tested: random forests with parametric bootstrap, weighted random forests, variable output smearing random forests, and weighted support vector regression. When using fit-quality metrics, at least one of these methods led to a statistically significant improvement on 31 of the 40 datasets. In the best case, these methods led to a 22% reduction in the root-mean-squared error of the models. Overall, our results demonstrate that by adapting data processing to account for curve fit quality, we can improve predictive performance across a range of different data sets.  \n■ INTRODUCTION  \nQuantitative structure activity relationship (QSAR) models are used to identify high-activity molecules for drug design and other bioactivity problems.1,2 Based on previously collected experimental data, models are constructed to predict the activity of untested molecules. While these models are often a useful tool in drug design processes,3−5 they are restricted by the quality of the available data.6,7 In this work, we focus on QSAR data sets where the relative quality of different data points varies significantly, and demonstrate that incorporating information about data point quality into the modeling process leads to improved predictions.  \nOur experiments used data sets containing dose−response experiments for different molecules. These experimental results are converted to EC50 values via a curve fitting procedure. An EC50 is the concentration at which a molecule induces a response of 50% of the maximum effect and is a commonly used metric in small molecule drug design.8 While EC50s are useful metrics, calculating reliable values from experimental data can be challenging.9 Figure 1 illustrates this challenge, showing examples of both a reliable (a) and unreliable (b) fit of EC50 curves.  \nThe standard recommendations are to ignore data points for which a reliable EC50 cannot be established. However, not only can this approach introduce bias, but it results in  \nexperimental data being ignored by the model. Systematic bias can occur because molecules with very low or very high EC50scan often end up with unreliable fits. For example, a common practice is to test all molecules at the same set of concentrations (as is the case in all of the PubChem datasets used in our study) and to discard calculated EC50 values that do not have two measurements above and below the EC50.10 This would result in excluding molecules with either very high or very low activity.  \nRather than discarding data with uncertain fits, we propose incorporating information about fit quality directly into the QSAR modeling process. During curve fitting, alongside the EC50, we calculate a “fit-quality metric”, a numerical score reflecting how closely t","cbCaihMRJgqYvXr3","https://ap.wps.com/l/cbCaihMRJgqYvXr3","pdf",2522365,1,12,"English","en",105,"# Abstract\n# Introduction\n## QSAR and EC50 modeling\n## Curve-fit reliability and bias from discarding uncertain EC50s\n## Fit-quality metric and uncertainty estimation\n# Methods and experimental setup\n## Datasets from PubChem and BASF\n## Machine learning models evaluated\n# Results and performance impact","[{\"question\":\"Why is curve fitting a limitation in early-stage machine learning drug design?\",\"answer\":\"Curve fitting reduces raw dose-response results to a single EC50 per molecule, which discards information about how reliable the fitted curve is. This lost fit-quality information limits the reliability of downstream model predictions.\"},{\"question\":\"How does the proposed fit-quality metric improve EC50-based modeling?\",\"answer\":\"During curve fitting, the method computes a numerical fit-quality score that reflects how closely the curve matches experimental points. Lower scores indicate a better fit and thus lower uncertainty in the EC50, enabling uncertainty-aware learning.\"},{\"question\":\"Which machine learning approaches were tested and how were they evaluated?\",\"answer\":\"The study tested four methods: random forests with parametric bootstrap, weighted random forests, variable output smearing random forests, and weighted support vector regression. Using the fit-quality metric, at least one method delivered statistically significant improvements on 31 of 40 datasets, with up to a 22% RMSE reduction.\"}]","Improved Machine Learning Predictions of EC50s Using Uncertainty Estimation from Dose-Response Data | PDF",1785815717,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"improved-machine-learning-predictions-of-ec50s-using-uncertainty-estimation-from-dose-response-data","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/improved-machine-learning-predictions-of-ec50s-using-uncertainty-estimation-from-dose-response-data/123279/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is curve fitting a limitation in early-stage machine learning drug design?","Question",{"text":75,"@type":76},"Curve fitting reduces raw dose-response results to a single EC50 per molecule, which discards information about how reliable the fitted curve is. This lost fit-quality information limits the reliability of downstream model predictions.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed fit-quality metric improve EC50-based modeling?",{"text":80,"@type":76},"During curve fitting, the method computes a numerical fit-quality score that reflects how closely the curve matches experimental points. Lower scores indicate a better fit and thus lower uncertainty in the EC50, enabling uncertainty-aware learning.",{"name":82,"@type":73,"acceptedAnswer":83},"Which machine learning approaches were tested and how were they evaluated?",{"text":84,"@type":76},"The study tested four methods: random forests with parametric bootstrap, weighted random forests, variable output smearing random forests, and weighted support vector regression. Using the fit-quality metric, at least one method delivered statistically significant improvements on 31 of 40 datasets, with up to a 22% RMSE reduction.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]