[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118106-en":3,"doc-seo-118106-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118106,5909877438554,"Maeve","https://ap-avatar.wpscdn.com/avatar/5600025385ad2bf12a7?_k=1778553567797529272",8,"Research & Report","Interpretable molecular encodings and representations for machine learning tasks","Molecular encodings and representations are central to machine learning breakthroughs in biomedical classification, especially for peptides and proteins. The work introduces Interpretable Carbon-based Array of Neighborhoods (iCAN), a structured counting-array encoding that captures carbon-atom neighborhoods and increases model usefulness without sacrificing interpretability. iCAN enables comparison of molecular neighborhoods, detection of repeating patterns, and relevance heat-map visualization for a given dataset. In a large peptide classification benchmark, iCAN surpasses its predecessor encoding, and on proteins it outperforms a lead-structure baseline on 71% of datasets.","Computational and Structural Biotechnology Journal 23 (2024) 2326–2336  \nContents lists available at ScienceDirect  \nComputational and Structural Biotechnology Journal  \njournal [homepage: www.elsevier.com/locate/csbj](homepage: www.elsevier.com/locate/csbj)  \n| Research Article\u003Cbr>Interpretable molecular encodings and representations for machine learning tasks\u003Cbr>Moritz Weckbecker a, 1 , Aleksandar Anžel a, 1 , Zewen Yang a, 1 , Georges Hattab a,b,∗\u003Cbr>a Center for Artiﬁcial Intelligence in Public Health Research,(ZKI-PH), Robert Koch Institute, Nordufer 20, Berlin, 13353, Berlin, Germany b Department of Mathematics and Computer science Freie Universität, Arnimallee 14, Berlin, 14195, Berlin, Germany |  |  |  |\n| --- | --- | --- | --- |\n| A R T I C L E I N F O |  | A B S T R A C T |  |\n| Dataset link: [https://github.com/ghattab/](https://github.com/ghattab/)[ ](https://github.com/ghattab/)[iCAN/tree/main/Data/Original_datasets](iCAN/tree/main/Data/Original_datasets) |  | Molecular encodings and their usage in machine learning models have demonstrated signiﬁcant breakthroughsin biomedical applications, particularly in the classiﬁcation of peptides and proteins. To this end, we propose anew encoding method: Interpretable Carbon-based Array of Neighborhoods (iCAN). Designed to address machine learning models’ need for more structured and less ﬂexible input, it captures the neighborhoods of carbon atoms in a counting array and improves the utility of the resulting encodings for machine learning models. The iCAN method provides interpretable molecular encodings and representations, enabling the comparison of molecular neighborhoods, identiﬁcation of repeating patterns, and visualization of relevance heat maps for a given data set. When reproducing a large biomedical peptide classiﬁcation study, it outperforms its predecessor encoding. When extended to proteins, it outperforms a lead structure-based encoding on 71% of the data sets. Our method oﬀers interpretable encodings that can be applied to all organic molecules, including exotic amino acids, cyclic peptides, and larger proteins, making it highly versatile across various domains and data sets. This work establishes a promising new direction for machine learning in peptide and protein classiﬁcation in biomedicine and healthcare, potentially accelerating advances in drug discovery and disease diagnosis. |  |\n| Keywords: Explainable Interpretable Molecular encoding Representation Machine learning |  |  |  |\n\n1. Introduction  \nMolecular ﬁngerprinting is a cornerstone of in silico molecular studies, virtual screening and machine learning (ML) applications in the ﬁeld of molecular sciences [1]. Speciﬁcally, a molecular ﬁngerprint encodes a molecule by converting its molecular structure into a bit string, which enables the application of various mathematical and computational methods to process, analyze and visually represent the molecule [2]. In the context of ML, encoding a molecule into a machinereadable format allows ML models to eﬀectively capture and learn its inherent chemical structure and properties.  \nDespite the widespread use of molecular ﬁngerprinting, there is a growing need for alternative encoding methods that can improve the performance of ML models in various molecular science applications. The current vast biomedical space and the limited existence of labeled and balanced data already pose challenges to most ML methods. Moreover, the lack of interpretable molecular encodings hinders the interpretability of models in biomedical research and applications. This poses signiﬁcant limitations that need to be addressed, particu-  \n* Corresponding author.  \n[E-mail addresses:](E-mail addresses: AnzelA@rki.de)[ AnzelA@rki.de](E-mail addresses: AnzelA@rki.de) (A. Anžel), [HattabG@rki.de](HattabG@rki.de) (G. Hattab).  \n1 These authors contributed equally to this work.  \nlarly in the development of novel interpretable molecular representations for ML, a concept that has not b","cbCaitgyiqYRVIAX","https://ap.wps.com/l/cbCaitgyiqYRVIAX","pdf",1086705,1,11,"English","en",105,"# Introduction\n## Molecular fingerprinting and the need for alternatives\n## Similar property principle and applications\n# Proposed method: iCAN\n## Interpretable carbon-based neighborhood encodings\n## Visualization and interpretability outputs\n# Experimental evaluation\n## Peptide classification benchmark\n## Protein extension and comparison","[{\"question\":\"What problem does iCAN address in molecular machine learning encodings?\",\"answer\":\"It targets the need for more structured, less flexible inputs and the lack of interpretable molecular encodings that limit understanding in biomedical research.\"},{\"question\":\"How does iCAN represent molecules for machine learning tasks?\",\"answer\":\"It captures carbon-atom neighborhoods using a counting array, producing molecular encodings that are interpretable and more useful to ML models.\"},{\"question\":\"How does iCAN perform compared with existing encodings?\",\"answer\":\"It outperforms its predecessor in a large biomedical peptide classification study, and when extended to proteins it outperforms a lead-structure-based encoding on 71% of evaluated datasets.\"}]","Interpretable molecular encodings and representations for machine learning tasks | PDF",1785681644,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"interpretable-molecular-encodings-and-representations-for-machine-learning-tasks","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/interpretable-molecular-encodings-and-representations-for-machine-learning-tasks/118106/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does iCAN address in molecular machine learning encodings?","Question",{"text":75,"@type":76},"It targets the need for more structured, less flexible inputs and the lack of interpretable molecular encodings that limit understanding in biomedical research.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does iCAN represent molecules for machine learning tasks?",{"text":80,"@type":76},"It captures carbon-atom neighborhoods using a counting array, producing molecular encodings that are interpretable and more useful to ML models.",{"name":82,"@type":73,"acceptedAnswer":83},"How does iCAN perform compared with existing encodings?",{"text":84,"@type":76},"It outperforms its predecessor in a large biomedical peptide classification study, and when extended to proteins it outperforms a lead-structure-based encoding on 71% of evaluated datasets.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]