[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123968-en":3,"doc-seo-123968-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123968,2336464648322,"Aria","https://ap-avatar.wpscdn.com/avatar/2200025388227c56fec?_k=1778556882303663488",8,"Research & Report","Machine learning-assisted search for novel coagulants - when machine learning can be efficient even if data availability is low","Design of new drugs requires candidate molecules to satisfy multiple criteria while minimizing side effects and avoiding off-target effects. Experimental data on molecular properties continues to grow, enabling data-driven methods, yet standard predictive machine learning can fail when target-specific datasets are scarce. This work proposes a deep-learning framework using an autoencoder representation of chemical space and sampling strategies to generate novel coagulant candidates, including validation on known anticoagulant targets and comparison with MegaMolBART.","arXiv :2401 .01811v1 [ q-bio .BM] 3 Jan 2024  \nMachine learning-assisted search for novel coagulants: when machine learning can be efficient even if data  \navailability is low  \nAndrij Rovenchak∗†, Maksym Druchok∗‡  \nJanuary 4, 2024  \nAbstract  \nDesign of new drugs is a challenging process: a candidate molecule should satisfy multiple conditions to act properly and make the least side-effect – perfect candidates selectively attach to and influence only targets, leaving off-targets intact. The amount of experimental data about various properties of molecules constantly grows, promoting data-driven approaches. However, the applicability of typical predictive machine learning techniques can be substantially limited by a lack of experimental data about a particular target. For example, there are many known Thrombin inhibitors (acting as anticoagulants), but a very limited number of known Protein C inhibitors (coagulants) . In this study, we present our approach to suggest new inhibitor candidates by building an effective representation of chemical space. For this aim, we developed a deep learning model – autoencoder, trained on a large set of molecules in the SMILES format to map the chemical space. Further, we applied different sampling strategies to generate novel coagulant candidates. Symmetrically, we tested our approach on anticoagulant candidates, where we were able to predict their inhibition towards Thrombin. We also compare our approach with MegaMolBART – another deep learning generative model, but exploiting similar principles of navigation in a chemical space.  \nKeywords: molecular design, machine learning, coagulants, anticoagulants.  \n∗ SoftServe, Inc. , 2d Sadova St. , 79021 Lviv, Ukraine  \n†Professor Ivan Vakarchuk Department for Theoretical Physics, Ivan Franko National University of Lviv, 12 Drahomanov St., 79005 Lviv, Ukraine  \n‡Institute for Condensed Matter Physics, 1 Svientsitskii St. , 79011 Lviv, Ukraine  \nThis study employs machine learning to generate new drugs, emphasizing cases with low data availability. Focusing on coagulants, underrepresented in databases, our approach generates molecular encodings based on the assumption that similar structures share properties. Strategies tested on anticoagulants are applied to discover novel coagulant candidates, navigating the encoding space.  \nINTRODUCTION  \nThis study is intended to explore a path of suggesting novel compounds by machine learning (ML) techniques with a specific case of limited data availability. In particular, we aim to find new coagulants, which are weakly represented in specialized databases, as much as, for example, anticoagulants. But before we dive into coagulants specifics, we would like to give a broader overview of current status of ML applications in drug design. A drug discovery process begins with localization of a disease cause, drafting of a list of drug candidates, and screening them in silico. During this screening, multiple properties of compounds are assessed to filter out least potent candidates and shrink the list before it gets to in vitro tests. Obviously, drafting and filtering influence the effectiveness of the downstream drug design stages quite substantially, as passing too many weak candidates will waste time and resources. Since structure and chemical composition of compounds define their properties, the task of establishing Quantitative Structure-Activity/Property Relationship (QSAR) relates structure of compounds with their physical/chemical properties, e.g., solubility in water or organic solvents, melting temperature, solvation energies, etc. 1–3  \nDuring the last decade, we witness a wave of ML-powered approaches in biochemical domain (see, for example, Refs. 4–10) . There are a few favoring factors here: (i) successes of ML in other areas, like computer vision, autonomous driving, natural language processing,(ii) growing availability of data, (iii) complex problems can be solved in a data-driven manner (in partic","cbCailXg2mx6FXLP","https://ap.wps.com/l/cbCailXg2mx6FXLP","pdf",2425127,1,40,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"Why can typical predictive machine learning be ineffective in drug or inhibitor discovery?\",\"answer\":\"Its effectiveness can be substantially limited when experimental data for a specific target is scarce, restricting training and generalization.\"},{\"question\":\"What is the core method proposed to search for novel coagulant candidates?\",\"answer\":\"The approach builds an effective representation of chemical space using a deep learning autoencoder trained on a large SMILES-format molecule set, then applies sampling strategies to generate candidates.\"},{\"question\":\"How is the approach evaluated and compared in the study?\",\"answer\":\"The model is tested on anticoagulant candidates by predicting inhibition toward Thrombin, and the method is compared with the generative model MegaMolBART, which uses similar chemical-space navigation principles.\"}]","Machine learning-assisted search for novel coagulants - when machine learning can be efficient even if data availability is low | PDF",1785819492,101,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"machine-learning-assisted-search-for-novel-coagulants-when-machine-learning-can-be-efficient-even-if-data-availability-is-low","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/machine-learning-assisted-search-for-novel-coagulants-when-machine-learning-can-be-efficient-even-if-data-availability-is-low/123968/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why can typical predictive machine learning be ineffective in drug or inhibitor discovery?","Question",{"text":75,"@type":76},"Its effectiveness can be substantially limited when experimental data for a specific target is scarce, restricting training and generalization.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the core method proposed to search for novel coagulant candidates?",{"text":80,"@type":76},"The approach builds an effective representation of chemical space using a deep learning autoencoder trained on a large SMILES-format molecule set, then applies sampling strategies to generate candidates.",{"name":82,"@type":73,"acceptedAnswer":83},"How is the approach evaluated and compared in the study?",{"text":84,"@type":76},"The model is tested on anticoagulant candidates by predicting inhibition toward Thrombin, and the method is compared with the generative model MegaMolBART, which uses similar chemical-space navigation principles.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":21,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]