[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121247-en":3,"doc-seo-121247-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},121247,13056703019404,"Miles","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Machine learning approaches for predicting protein-ligand binding sites from sequence data","Proteins contain specific protein-ligand binding sites that enable essential molecular interactions in biochemical reactions. Accurate identification is crucial for computational drug discovery, since it helps reveal therapeutic targets and supports treatment development. This mini review summarizes studies applying machine learning to predict binding sites from sequence data, emphasizing recent advances in embedding methods and learning architectures. It also highlights technical challenges, active debates, literature gaps, and future directions for improving sequence-based prediction.","TYPE Mini Review  \nPUBLISHED 03 February 2025 DOI 10.3389/fbinf.2025.1520382  \nOPEN ACCESS  \nEDITED BY  \nWen Wei,  \nArizona State University, United States  \nREVIEWED BY  \nKumar Yugandhar,  \nCornell University, United States Minh Nguyen,  \nBioinformatics Institute (A∗STAR), Singapore  \n*CORRESPONDENCE  \nOrhun Vural,  \n [orhun@uab.edu](orhun@uab.edu)  \nRECEIVED 31 October 2024  \nACCEPTED 10 January 2025  \nPUBLISHED 03 February 2025  \nCITATION  \nVural O and Jololian L (2025) Machine learning approaches for predicting protein-ligand binding sites from sequence data.  \nFront. Bioinform. 5:1520382 .  \ndoi: 10.3389/fbinf.2025.1520382  \nCOPYRIGHT  \n© 2025 Vural and Jololian. This is an  \nopen-access article distributed under the terms of the Creative Commons Attribution License (CC BY) . The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.  \nMachine learning approaches for predicting protein-ligand binding sites from sequence data  \nOrhun Vural* and Leon Jololian  \nDepartment of Electrical and Computer Engineering, The University of Alabama at Birmingham, Birmingham, AL, United States  \nProteins, composed of amino acids, are crucial for a wide range of biological functions. Proteins have various interaction sites, one of which is the protein-ligand binding site, essential for molecular interactions and biochemical reactions. These sites enable proteins to bind with other molecules, facilitating key biological functions. Accurate prediction of these binding sites is pivotal in computational drug discovery, helping to identify therapeutic targets and facilitate treatment development. Machine learning has made significant contributions to this field by improving the prediction of proteinligand interactions. This paper reviews studies that use machine learning to predict protein-ligand binding sites from sequence data, focusing on recent advancements. The review examines various embedding methods and machine learning architectures, addressing current challenges and the ongoing debatesin the field. Additionally, research gaps in the existing literature are highlighted, and potential future directions for advancing the field are discussed. This study provides a thorough overview of sequence-based approaches for predicting protein-ligand binding sites, offering insights into the current state of research and future possibilities.  \nKEYWORDS  \nprotein-ligand binding sites, computational drug discovery, sequence-based methods, deep learning, binding prediction  \n1 Introduction  \nProtein-ligand binding sites are specific regions on proteins where various ligands—including small organic molecules, peptides, nucleotides, and proteins—can attach or bind (Zhao et al., 2020) . Although experimental laboratory methods identify these regions with the highest accuracy, they are generally costly and time-consuming (Sadybekov and Katritch, 2023) . Therefore, computational approaches to drug discovery have become increasingly important. These computational methods offer distinct advantages by reducing costs and speeding up identifying and optimizing potential drug candidates (Gupta et al., 2021). Predicting protein-ligand binding sites is a critical component of computational drug discovery, essential for pinpointing viable drug targets and advancing the development of new therapeutics (Stank et al., 2016) . Recent advancements in machine learning have significantly improved this field by introducing sophisticated computational techniques to analyze the complex interactions between proteins and ligands (Xia et al., 2024) . While traditional methods based on geometry, energy, or templates have been successful, deep learning has recently achieved much better results (Gagliardi ","cbCaipF1Kc1igwlQ","https://ap.wps.com/l/cbCaipF1Kc1igwlQ","pdf",2348495,1,9,"English","en",105,"# Introduction\n## Sequence-based vs structure-based prediction\n## Background on binding sites and drug discovery\n## Role of deep learning and architectures","[{\"question\":\"Why are protein-ligand binding sites important for computational drug discovery?\",\"answer\":\"They are specific protein regions where ligands bind, enabling molecular interactions and biochemical reactions. Predicting them supports identifying viable drug targets and advancing therapeutics development.\"},{\"question\":\"What major input categories are used for predicting binding sites?\",\"answer\":\"Models are commonly divided into structure-based and sequence-based approaches, depending on whether they use spatial/structural information or sequence data.\"},{\"question\":\"How do deep learning methods improve prediction compared with traditional approaches?\",\"answer\":\"Deep learning can learn complex patterns directly from raw data and generalize better across diverse datasets, often achieving stronger results than geometry-, energy-, or template-based methods.\"}]","Machine learning approaches for predicting protein-ligand binding sites from sequence data | PDF",1785734587,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"machine-learning-approaches-for-predicting-protein-ligand-binding-sites-from-sequence-data","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/machine-learning-approaches-for-predicting-protein-ligand-binding-sites-from-sequence-data/121247/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why are protein-ligand binding sites important for computational drug discovery?","Question",{"text":75,"@type":76},"They are specific protein regions where ligands bind, enabling molecular interactions and biochemical reactions. Predicting them supports identifying viable drug targets and advancing therapeutics development.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What major input categories are used for predicting binding sites?",{"text":80,"@type":76},"Models are commonly divided into structure-based and sequence-based approaches, depending on whether they use spatial/structural information or sequence data.",{"name":82,"@type":73,"acceptedAnswer":83},"How do deep learning methods improve prediction compared with traditional approaches?",{"text":84,"@type":76},"Deep learning can learn complex patterns directly from raw data and generalize better across diverse datasets, often achieving stronger results than geometry-, energy-, or template-based methods.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]