[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122551-en":3,"doc-seo-122551-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122551,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","Can AI-Predicted Complexes Teach Machine Learning to Compute Drug Binding Affinity?","This study evaluates the feasibility of using co-folding models to generate synthetic data for training machine learning-based scoring functions (MLSFs) that predict drug binding affinity. Results show that improvements depend on the structural quality of augmented complexes. The work proposes simple heuristics to identify high-quality co-folding predictions without requiring reference experimental structures, allowing selected predictions to substitute for experimental inputs during MLSF training. Findings guide future data augmentation designs based on cofolding models.","This article is licensed under CC-BY 4.0   \n[pubs.acs.org/jcim](pubs.acs.org/jcim)  Letter   \nCan AI-Predicted Complexes Teach Machine Learning to Compute Drug Binding Affinity?  \nWei-Tse Hsu, Savva Grevtsev, Anna M. Herz, Thomas Douglas, Aniket Magarkar,* and Philip C. Biggin *  \n Cite This: J. Chem. Inf. Model. 2025, 65, 13051−13056  \nRead Online  \n\n|  |  |  |  |  |  |\n| --- | --- | --- | --- | --- | --- |\n| ACCESS   | Metrics & More |  |  Article Recommendations |  | *sı Supporting Information |\n\nABSTRACT: We evaluate the feasibility of using co-folding models for synthetic data augmentation in training machine learning-based scoring functions (MLSFs) for binding affinity prediction. Our results show that performance gains depend critically on the structural quality of augmented data. In light of this, we established simple heuristics for identifying high-quality co-folding predictions without reference structures, enabling them to substitute for experimental structures in MLSF training. Our study informs future data augmentation strategies based on cofolding models.  \n■ INTRODUCTION  \nOver the past decades, machine learning-based scoring functions (MLSFs) have gained increasing popularity in computer-aided drug discovery.1 By leveraging 3D structures of binding complexes􀀁usually protein-ligand binding complexes􀀁these models predict binding affinities in a fraction of the time required by physics-based simulation methods such as alchemical free energy perturbation,2 while achieving arguably comparable accuracy in some scenarios.3 During training, they often rely on experimental structures of binding complexes, representing binding interfaces with underlying architectures ranging from feed-forward neural networks,4 convolutional neural networks (CNNs),5 transformers,6 to graph neural networks (GNNs).3,7 However, the data of high-resolution experimental complexes with matched binding affinity measurements remain rare, limiting both the scale and diversity of training data sets available for these models.  \nTo address this scarcity, several efforts have emerged to synthetically augment training data sets using computational modeling. One notable example is BindingNet,8 which uses protein structures from PDBbind9 as templates and models new complexes by aligning structurally similar ChEMBL 10 ligands to the reference ligands based on their maximum common substructures. With this template-based modeling approach, BindingNet v1 generated approximately 70K protein-ligand complexes with associated activity data from ChEMBL. Its successor, BindingNet v2, 11 introduced a hierarchical variation of the modeling pipeline to accommodate less similar candidate ligands, further expanding the data set to roughly 700 K complexes. Recent studies have demonstrated that the inclusion of BindingNet v1 improves the performance of MLSFs,3 though BindingNet v2 has so far only been used to train the docking model Uni-Mol, 12 where  \nimproved success rates in PoseBusters 13 sanity checks were observed, i.e. more physical binding poses were generated.  \nOne inherent drawback of these template-based modeling approaches, however, is their reliance on high-quality, experimentally determined protein structures as templates, which restricts the extent of data augmentation. Moreover, these methods implicitly assume that structurally similar ligands that bind to the same protein receptor share the same binding mode, which does not always hold in practice. Recent advances in co-folding models, such as AlphaFold3 (AF3), 14 Chai-1,15 and Boltz, 16, 17 enable de novo structure prediction of protein-ligand complex structures, offering a promising alternative to further expand the scope of binding  \ncomplex data using Boltz-1x  \nsets. Indeed, a large-scale data set generated has been recently proposed by Lemos et al.18  \nYet, the use of co-folding predictions for large-scale data set generation has not been systematically examined in the context of training MLSFs.","cbCaipSqhTs8YcEM","https://ap.wps.com/l/cbCaipSqhTs8YcEM","pdf",1826191,1,6,"English","en",105,"# Abstract\n# Introduction\n## Machine learning scoring functions in drug discovery\n## Template-based synthetic augmentation (BindingNet)\n## Co-folding models for de novo complex prediction\n## Study design and research questions","[{\"question\":\"How does synthetic data augmentation affect MLSF performance in this study?\",\"answer\":\"Performance gains from augmentation depend critically on the structural quality of the generated synthetic complexes, rather than on augmentation alone.\"},{\"question\":\"Can co-folding predictions replace experimental structures for MLSF training?\",\"answer\":\"The study examines whether co-folding outputs can substitute for or complement experimental structures, and develops methods to select high-quality predictions for training.\"},{\"question\":\"How are high-quality co-folding predictions identified without reference structures?\",\"answer\":\"The authors establish simple heuristics that filter co-folding predictions based on their quality, enabling useful synthetic examples even when reference structures are unavailable.\"}]","Can AI-Predicted Complexes Teach Machine Learning to Compute Drug Binding Affinity? | PDF",1785811238,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"can-ai-predicted-complexes-teach-machine-learning-to-compute-drug-binding-affinity","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/can-ai-predicted-complexes-teach-machine-learning-to-compute-drug-binding-affinity/122551/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How does synthetic data augmentation affect MLSF performance in this study?","Question",{"text":75,"@type":76},"Performance gains from augmentation depend critically on the structural quality of the generated synthetic complexes, rather than on augmentation alone.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Can co-folding predictions replace experimental structures for MLSF training?",{"text":80,"@type":76},"The study examines whether co-folding outputs can substitute for or complement experimental structures, and develops methods to select high-quality predictions for training.",{"name":82,"@type":73,"acceptedAnswer":83},"How are high-quality co-folding predictions identified without reference structures?",{"text":84,"@type":76},"The authors establish simple heuristics that filter co-folding predictions based on their quality, enabling useful synthetic examples even when reference structures are unavailable.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]