[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-124845-en":3,"doc-seo-124845-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},124845,4398048949847,"Eliana","https://ap-avatar.wpscdn.com/avatar/400002536579ef2da7f?_k=1778318612642679267",8,"Research & Report","A Small Step Toward Generalizability - Training a Machine Learning Scoring Function for Structure-Based Virtual Screening","Machine learning scoring functions for small-molecule–protein binding aim to approximate interaction energies, yet many learn dataset biases and generalize only to similar targets. This work builds a bias-robust, structure-based virtual screening scoring function using rigorous train/test filtering and evaluates it on the CASF-2016 benchmark with performance comparable to leading methods. Attribution analysis identifies binding interactions linked to important bonds and correlates them with distance-based interaction profiles. The extracted pharmacophores support fragment elaboration, improving docking scores and demonstrating a deep learning approach for extracting structural information for molecule design.","[pubs.acs.org/jcim](pubs.acs.org/jcim)  Article   \nA Small Step Toward Generalizability: Training a Machine Learning Scoring Function for Structure-Based Virtual Screening  \nJack Scantlebury,† Lucy Vost,† Anna Carbery, Thomas E. Hadfield, Oliver M. Turnbull, Nathan Brown, Vijil Chenthamarakshan, Payel Das, Harold Grosjean, Frank von Delft, and Charlotte M. Deane*  \n Cite This: J. Chem. Inf. Model. 2023, 63, 2960−2974  \nRead Online  \nDownloaded via 193.237.195.116 on May 22, 2023 at 14:07:07 (UTC) . See [https://pubs.acs.org/sharingguidelines](https://pubs.acs.org/sharingguidelines) for options on how to legitimately share published articles.  \nACCESS  \n Metrics & More  \n Article Recommendations  \n*sı   \nSupporting Information  \nABSTRACT: Over the past few years, many machine learning-based scoring functions for predicting the binding of small molecules to proteins have been developed. Their objective is to approximate the distribution which takes two molecules as input and outputs the energy of their interaction. Only a scoring function that accounts for the interatomic interactions involved in binding can accurately predict binding affinity on unseen molecules. However, many scoring functions make predictions based on data set biases rather than an understanding of the physics of binding. These scoring functions perform well when tested on similar targets to those in the training set but fail to generalize to dissimilar targets. To test what a machine learning-based scoring function has learned, input attribution, a technique for learning which features are important to a model when making a prediction on a particular data point, can be applied. If a model successfully learns something beyond data set biases, attribution should give insight into the important binding interactions that are taking place. We built a machine learning-based scoring function that aimed to avoid the influence of bias via thorough train and test data set filtering and show that it achieves comparable performance on the Comparative Assessment of Scoring Functions, 2016 (CASF-2016) benchmark to other leading methods. We then use the CASF-2016 test set to perform attribution and find that the bonds identified as important by PointVS, unlike those extracted from other scoring functions, have a high correlation with those found by a distance-based interaction profiler. We then show that attribution can be used to extract important binding pharmacophores from a given protein target when supplied with a number of bound structures. We use this information to perform fragment elaboration and see improvements in docking scores compared to using structural information from a traditional, data-based approach. This not only provides definitive proof that the scoring function has learned to identify some important binding interactions but also constitutes the first deep learning-based method for extracting structural information from a target for molecule design.  \n■ INTRODUCTION  \nThe recent explosion of machine learning (ML) across all scientific disciplines has been accompanied by concerns regarding the generalizability of methods used. Various studies have found data leakage to be present in ML applications, resulting in overly optimistic performances being reported for the tools in question. 1−4 The field of drug discovery is by no means exempt from this: models that learn unintended features from training data sets are extremely sensitive to small data set distribution shifts and cannot make reliable predictions on out-of-distribution data points. Practically, this renders them incapable of assisting in the development of drugs for novel targets.5−7  \nTypically led by human experts, drug development is a time-consuming and expensive process, with recent estimates placing the time required to reach clinical trials at 8.3 years, and the median cost at $985 million.8 Computational techniques offer a promising alternate route to human-led  \ndesign. One such comp","cbCaitybI3Okxuwx","https://ap.wps.com/l/cbCaitybI3Okxuwx","pdf",2500003,1,15,"English","en",105,"# Abstract\n## Generalization challenge and dataset bias\n## Method: bias mitigation and scoring function training\n## Evaluation on CASF-2016\n## Attribution for binding interaction interpretation\n## Pharmacophore extraction and fragment elaboration","[{\"question\":\"Why do some machine learning scoring functions fail to generalize to new protein targets?\",\"answer\":\"Many models exploit dataset biases rather than physically grounded interatomic interactions, leading to unreliable predictions under distribution shifts.\"},{\"question\":\"How was the proposed scoring function designed to reduce bias?\",\"answer\":\"It uses thorough train and test data set filtering to minimize the influence of dataset biases before benchmarking performance.\"},{\"question\":\"How can attribution be used in this work and what does it help reveal?\",\"answer\":\"Attribution identifies features important to predictions on test data points, and the resulting important bonds correlate with interactions found by a distance-based interaction profiler; these insights guide pharmacophore extraction and fragment elaboration.\"}]","A Small Step Toward Generalizability - Training a Machine Learning Scoring Function for Structure-Based Virtual Screening | PDF",1785894972,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"a-small-step-toward-generalizability-training-a-machine-learning-scoring-function-for-structure-based-virtual-screening","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/a-small-step-toward-generalizability-training-a-machine-learning-scoring-function-for-structure-based-virtual-screening/124845/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do some machine learning scoring functions fail to generalize to new protein targets?","Question",{"text":75,"@type":76},"Many models exploit dataset biases rather than physically grounded interatomic interactions, leading to unreliable predictions under distribution shifts.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How was the proposed scoring function designed to reduce bias?",{"text":80,"@type":76},"It uses thorough train and test data set filtering to minimize the influence of dataset biases before benchmarking performance.",{"name":82,"@type":73,"acceptedAnswer":83},"How can attribution be used in this work and what does it help reveal?",{"text":84,"@type":76},"Attribution identifies features important to predictions on test data points, and the resulting important bonds correlate with interactions found by a distance-based interaction profiler; these insights guide pharmacophore extraction and fragment elaboration.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]