[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120138-en":3,"doc-seo-120138-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120138,8796095461564,"Liam","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","On Machine Learning Approaches for Protein-Ligand Binding Affinity Prediction","Binding affinity optimization is pivotal in early-stage drug discovery, yet the relative performance of existing machine learning methods for ligand potency prediction is not fully understood. This study benchmarks classical tree-based models and advanced neural networks for protein-ligand binding affinity prediction, covering ligand-only 2D RDKit embeddings and Large Language Model (LLM) ligand representations, plus 3D neural networks that use bound protein-ligand conformations. Experiments span multiple standard datasets and settings including classification, ranking, regression, and active learning.","arXiv :2407 . 19073v1 [ q-bio .BM] 15 Jul 2024  \nOn Machine Learning Approaches for Protein-Ligand Binding Affinity Prediction  \nNikolai Schapina,b , Carles Navarroa,b , Albert Boub , and Gianni De Fabritiis∗ a,b,c  \naAcellera Labs, C/ Doctor Trueta 183, 08005 Barcelona, Spain b Computational Science Laboratory, Universitat Pompeu Fabra, PRBB, C/  \nDoctor Aiguader 88, 08003 Barcelona, Spain  \nc Instituci´o Catalana de Recerca i Estudis Avan¸cats (ICREA), Passeig Llu´ıs  \nCompanys 23, 08010 Barcelona, Spain  \nJuly 30, 2024  \nAbstract  \nBinding affinity optimization is crucial in early-stage drug discovery. While numerous machine learning methods exist for predicting ligand potency, their comparative efficacy remains unclear. This study evaluates the performance of classical tree-based models and advanced neural networks in protein-ligand binding affinity prediction. Our comprehensive benchmarking encompasses 2D models utilizing ligand-only RDKit embeddings and Large Language Model (LLM) ligand representations, as well as 3D neural networks incorporating bound protein-ligand conformations. We assess these models across multiple standard datasets, examining various predictive scenarios including classification, ranking, regression, and active learning. Results indicate that simpler models can surpass more complex ones in specific tasks, while 3D models leveraging structural information become increasingly competitive with larger training datasets containing compounds with labelled affinity data against multiple targets. Pre-trained 3D models, by incorporating protein pocket environments, demonstrate significant advantages in data-scarce scenarios for specific binding pockets. Additionally, LLM pretraining on 2D ligand data enhances complex model performance, providing versatile embeddings that outperform traditional RDKit features in computational efficiency. Finally, we show that combining 2D and 3D model strengths improves active learning outcomes beyond current state-of-the-art approaches. These findings offer valuable insights for optimizing machine learning strategies in drug discovery pipelines.  \nKeywords—binding affinity prediction, machine learning, tree-based models, graph neural networks, classification, regression, active learning, large language models  \n1 Introduction  \nBinding affinity is a crucial factor in the preclinical drug discovery of small-molecule therapeutics [1] . Traditionally assessed through experimental techniques, considerable advancements have been made in computational methods to predict protein-ligand binding affinities, including machine learned models. Nowadays, a wide range of tested model architectures exist that are employed in various predictive scenarios.  \nGeneral affinity prediction models are often trained on heterogeneous binding affinity datasets, employing different input types and architectures. For instance, ligand SMILES embeddings from Large Language Models (LLMs) can be used with or without protein sequence data, as seen in ChemBoost [2], DeepFusionDTA [3], AttentionDTA [4],  \nand DeepDTA [5] . Simplified models like XGBoost [6, 7] or one-dimensional convolutional neural networks can then map these embeddings to affinity values. Alternatively, chemical descriptors like QM energy terms [8, 9] or physicochemical descriptors from RDKit [10, 11, 12] are used with simpler models such as tree-based methods, support vector machines [13], Bayesian models, or neural networks. These can be further augmented with protein-ligand interaction fingerprints [14, 15, 16, 17, 18, 19, 20, 21, 22, 23], capturing ligand-target interactions as 2D vectors. 3D bound conformations can also be embedded using molecular graphs and graph neural networks [24, 25, 26], or represented through voxels, as done in KDeep [27] and other models [28, 29, 30, 31] . All these models can rank and prioritize compounds during virtual screening.  \nAmong the more precise computational methods are the free energy perturbatio","cbCairj7Cz2smJrm","https://ap.wps.com/l/cbCairj7Cz2smJrm","pdf",5740200,1,49,"English","en",105,"# Introduction\n## General affinity prediction models\n## More precise computational methods and ML alternatives\n## Classification models for virtual screening\n## Pretraining strategies for chemical neural networks","[{\"question\":\"What modeling approaches are benchmarked for protein-ligand binding affinity prediction?\",\"answer\":\"The document benchmarks classical tree-based models and advanced neural networks, including 2D ligand-only models (RDKit embeddings and LLM ligand representations) and 3D models using bound protein-ligand conformations.\"},{\"question\":\"What evaluation scenarios are used in the benchmarking study?\",\"answer\":\"The study evaluates models across standard datasets using classification, ranking, regression, and active learning settings.\"},{\"question\":\"What are the key findings regarding model complexity and data size?\",\"answer\":\"Simpler models can outperform more complex ones in specific tasks, while 3D structural models become increasingly competitive when trained on larger datasets. Pre-trained 3D models also show advantages in data-scarce scenarios for particular binding pockets.\"}]","On Machine Learning Approaches for Protein-Ligand Binding Affinity Prediction | PDF",1785728399,123,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"on-machine-learning-approaches-for-protein-ligand-binding-affinity-prediction","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/on-machine-learning-approaches-for-protein-ligand-binding-affinity-prediction/120138/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What modeling approaches are benchmarked for protein-ligand binding affinity prediction?","Question",{"text":75,"@type":76},"The document benchmarks classical tree-based models and advanced neural networks, including 2D ligand-only models (RDKit embeddings and LLM ligand representations) and 3D models using bound protein-ligand conformations.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What evaluation scenarios are used in the benchmarking study?",{"text":80,"@type":76},"The study evaluates models across standard datasets using classification, ranking, regression, and active learning settings.",{"name":82,"@type":73,"acceptedAnswer":83},"What are the key findings regarding model complexity and data size?",{"text":84,"@type":76},"Simpler models can outperform more complex ones in specific tasks, while 3D structural models become increasingly competitive when trained on larger datasets. Pre-trained 3D models also show advantages in data-scarce scenarios for particular binding pockets.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]