[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118681-en":3,"doc-seo-118681-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118681,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","Interpretable molecular encodings and representations for machine learning tasks - Research article","Molecular encodings play a key role in biomedical machine learning, especially for classifying peptides and proteins. An interpretable encoding approach, Interpretable Carbon-based Array of Neighborhoods (iCAN), is introduced to structure model inputs by counting carbon-atom neighborhoods. The method enables comparison of neighborhoods, detection of repeating patterns, and relevance heat-map visualization. On a large peptide classification study it exceeds a prior encoding, and on proteins it outperforms a lead-structure encoding on 71% of datasets, providing versatile interpretability across organic molecules for drug discovery and disease diagnosis.","Computational and Structural Biotechnology Journal 23 (2024) 2326–2336  \nContents lists available at ScienceDirect  \nComputational and Structural Biotechnology Journal  \njournal [homepage: www.elsevier.com/locate/csbj](homepage: www.elsevier.com/locate/csbj)  \n| Research Article\u003Cbr>Interpretable molecular encodings and representations for machine learning tasks\u003Cbr>Moritz Weckbecker a, 1 , Aleksandar Anžel a, 1 , Zewen Yang a, 1 , Georges Hattab a,b,∗\u003Cbr>a Center for Artiﬁcial Intelligence in Public Health Research,(ZKI-PH), Robert Koch Institute, Nordufer 20, Berlin, 13353, Berlin, Germany b Department of Mathematics and Computer science Freie Universität, Arnimallee 14, Berlin, 14195, Berlin, Germany |  |  |  |\n| --- | --- | --- | --- |\n| A R T I C L E I N F O |  | A B S T R A C T |  |\n| Dataset link: [https://github.com/ghattab/](https://github.com/ghattab/)[ ](https://github.com/ghattab/)[iCAN/tree/main/Data/Original_datasets](iCAN/tree/main/Data/Original_datasets) |  | Molecular encodings and their usage in machine learning models have demonstrated signiﬁcant breakthroughsin biomedical applications, particularly in the classiﬁcation of peptides and proteins. To this end, we propose anew encoding method: Interpretable Carbon-based Array of Neighborhoods (iCAN). Designed to address machine learning models’ need for more structured and less ﬂexible input, it captures the neighborhoods of carbon atoms in a counting array and improves the utility of the resulting encodings for machine learning models. The iCAN method provides interpretable molecular encodings and representations, enabling the comparison of molecular neighborhoods, identiﬁcation of repeating patterns, and visualization of relevance heat maps for a given data set. When reproducing a large biomedical peptide classiﬁcation study, it outperforms its predecessor encoding. When extended to proteins, it outperforms a lead structure-based encoding on 71% of the data sets. Our method oﬀers interpretable encodings that can be applied to all organic molecules, including exotic amino acids, cyclic peptides, and larger proteins, making it highly versatile across various domains and data sets. This work establishes a promising new direction for machine learning in peptide and protein classiﬁcation in biomedicine and healthcare, potentially accelerating advances in drug discovery and disease diagnosis. |  |\n| Keywords: Explainable Interpretable Molecular encoding Representation Machine learning |  |  |  |\n\n1. Introduction  \nMolecular ﬁngerprinting is a cornerstone of in silico molecular studies, virtual screening and machine learning (ML) applications in the ﬁeld of molecular sciences [1]. Speciﬁcally, a molecular ﬁngerprint encodes a molecule by converting its molecular structure into a bit string, which enables the application of various mathematical and computational methods to process, analyze and visually represent the molecule [2]. In the context of ML, encoding a molecule into a machinereadable format allows ML models to eﬀectively capture and learn its inherent chemical structure and properties.  \nDespite the widespread use of molecular ﬁngerprinting, there is a growing need for alternative encoding methods that can improve the performance of ML models in various molecular science applications. The current vast biomedical space and the limited existence of labeled and balanced data already pose challenges to most ML methods. Moreover, the lack of interpretable molecular encodings hinders the interpretability of models in biomedical research and applications. This poses signiﬁcant limitations that need to be addressed, particu-  \n* Corresponding author.  \n[E-mail addresses:](E-mail addresses: AnzelA@rki.de)[ AnzelA@rki.de](E-mail addresses: AnzelA@rki.de) (A. Anžel), [HattabG@rki.de](HattabG@rki.de) (G. Hattab).  \n1 These authors contributed equally to this work.  \nlarly in the development of novel interpretable molecular representations for ML, a concept that has not b","cbCaiu9DicGbc4Jp","https://ap.wps.com/l/cbCaiu9DicGbc4Jp","pdf",1082325,1,11,"English","en",105,"# Introduction\n## Molecular fingerprinting and need for interpretable encodings\n## Similar property principle and biomedical applications","[{\"question\":\"What is iCAN and what problem does it address?\",\"answer\":\"iCAN is an interpretable carbon-based encoding method that captures carbon-atom neighborhoods in a counting array. It is designed to provide more structured, less flexible inputs while improving interpretability for machine learning models in biomedical tasks.\"},{\"question\":\"How does iCAN support interpretability in molecular learning tasks?\",\"answer\":\"iCAN enables direct comparison of molecular neighborhoods, identification of repeating patterns, and visualization through relevance heat maps for a given dataset.\"},{\"question\":\"How does iCAN perform compared with prior encodings?\",\"answer\":\"When reproducing a large biomedical peptide classification study, iCAN outperforms its predecessor encoding. Extended to proteins, it outperforms a lead-structure-based encoding on 71% of the evaluated datasets.\"}]","Interpretable molecular encodings and representations for machine learning tasks - Research article | PDF",1785684862,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"interpretable-molecular-encodings-and-representations-for-machine-learning-tasks-research-article","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/interpretable-molecular-encodings-and-representations-for-machine-learning-tasks-research-article/118681/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04","2026-08-02",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What is iCAN and what problem does it address?","Question",{"text":76,"@type":77},"iCAN is an interpretable carbon-based encoding method that captures carbon-atom neighborhoods in a counting array. It is designed to provide more structured, less flexible inputs while improving interpretability for machine learning models in biomedical tasks.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does iCAN support interpretability in molecular learning tasks?",{"text":81,"@type":77},"iCAN enables direct comparison of molecular neighborhoods, identification of repeating patterns, and visualization through relevance heat maps for a given dataset.",{"name":83,"@type":74,"acceptedAnswer":84},"How does iCAN perform compared with prior encodings?",{"text":85,"@type":77},"When reproducing a large biomedical peptide classification study, iCAN outperforms its predecessor encoding. Extended to proteins, it outperforms a lead-structure-based encoding on 71% of the evaluated datasets.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]