[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118098-en":3,"doc-seo-118098-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118098,4398048950312,"Violet","https://ap-avatar.wpscdn.com/avatar/400002538284de19e3c?_k=1778320343897328908",8,"Research & Report","Exploring machine learning for untargeted metabolomics using molecular fingerprints","Metabolomics studies substrates and products of cellular metabolism and can inform preventive healthcare and pharmaceutical research, but large dataset analysis remains difficult due to limited and incompletely annotated pathway knowledge. A machine-learning framework inspired by drug discovery is presented, using molecular metabolite fingerprints to relate structure to experimental responses beyond known pathways. The method is evaluated on representational effectiveness while addressing class imbalance, data sparsity, high dimensionality, duplicate structural encoding, and interpretability, then applies feature-importance analysis to identify key chemical configurations. Results on Ataxia Telangiectasia and low-oxygen endothelial-cell datasets show effective prediction, consistency with known pathways, and discovery of additional affected metabolite groups for further study.","Computer Methods and Programs in Biomedicine 250 (2024) 108163  \nContents lists available at ScienceDirect  \nComputer Methods and Programs in Biomedicine  \njournal [homepage: www.elsevier.com/locate/cmpb](homepage: www.elsevier.com/locate/cmpb)  \n| Exploring machine learning for untargeted metabolomics using molecular ﬁngerprints |  |  |  |\n| --- | --- | --- | --- |\n| Christel Sirocchi a,∗, 1 , Federica Biancuccib, 1 , Matteo Donati a, Alessandro Bogliolo a, Mauro Magnanib, Michele Menottab, Sara Montagna a\u003Cbr>a Department of Pure and Applied Sciences, University of Urbino, Piazza della Repubblica, 13, Urbino, 61029, Italy b Department of Biomolecular Sciences, University of Urbino, Via Saﬃ 2, Urbino, 61029, Italy |  |  |  |\n| A R T I C L E I N F O |  | A B S T R A C T |  |\n| Keywords:\u003Cbr>Ataxia telangiectasia Mass spectrometry Molecular ﬁngerprinting Untargeted metabolomics Machine learning |  | Background: Metabolomics, the study of substrates and products of cellular metabolism, oﬀers valuable insights into an organism’s state under speciﬁc conditions and has the potential to revolutionise preventive healthcare and pharmaceutical research. However, analysing large metabolomics datasets remains challenging, with available methods relying on limited and incompletely annotated metabolic pathways.\u003Cbr>Methods: This study, inspired by well-established methods in drug discovery, employs machine learning on metabolite ﬁngerprints to explore the relationship of their structure with responses in experimental conditions beyond known pathways, shedding light on metabolic processes. It evaluates ﬁngerprinting eﬀectiveness in representing metabolites, addressing challenges like class imbalance, data sparsity, high dimensionality, duplicate structural encoding, and interpretable features. Feature importance analysis is then applied to reveal key chemical conﬁgurations aﬀecting classiﬁcation, identifying related metabolite groups.\u003Cbr>Results: The approach is tested on two datasets: one on Ataxia Telangiectasia and another on endothelial cells under low oxygen. Machine learning on molecular ﬁngerprints predicts metabolite responses eﬀectively, and feature importance analysis aligns with known metabolic pathways, unveiling new aﬀected metabolite groups for further study.\u003Cbr>Conclusion: In conclusion, the presented approach leverages the strengths of drug discovery to address critical issues in metabolomics research and aims to bridge the gap between these two disciplines. This work lays the foundation for future research in this direction, possibly exploring alternative structural encodings and machine learning models. |  |\n\n1. Introduction  \nMetabolomics is the study of small molecule substrates and products of cellular metabolism and provides valuable insights into the state of an organism under speciﬁc conditions [1]. Metabolomic proﬁling of diseased and healthy tissues is instrumental in discovering distinctive metabolic signatures and biomarkers, thereby aiding the development of screening tests and identiﬁcation of potential drug targets [2]. Additionally, metabolomics can help to assess the eﬀects of candidate treatments, evaluating the response at the metabolic level [3]. Therefore, metabolomics serves as an indispensable tool in preventive healthcare as well as pharmaceutical research and development, with the potential to enable timely disease detection and facilitate drug testing [4].  \n* Corresponding author.  \nE-mail address: [c.sirocchi2@campus.uniurb.it](c.sirocchi2@campus.uniurb.it) (C. Sirocchi).  \n1 These authors contributed equally to this work.  \nBeyond its clinical applications, metabolomics shows signiﬁcant potential in basic research for unravelling the mechanisms of action behind diseases and treatments. In this context, one of the prominent methods for analysing metabolomic data is pathway enrichment analysis, which identiﬁes metabolic pathways with a higher-than-expected abundance of aﬀected metabolites, oﬀering ins","cbCaiqTYEN1lnrFy","https://ap.wps.com/l/cbCaiqTYEN1lnrFy","pdf",1685081,1,15,"English","en",105,"# Introduction\n## Metabolomics background and clinical relevance\n## Pathway enrichment analysis and its limitations\n# Methods\n## Molecular fingerprint encoding and machine-learning modeling\n## Evaluation challenges and feature-importance interpretation\n# Results\n## Ataxia Telangiectasia dataset experiments\n## Low-oxygen endothelial cells dataset experiments\n# Conclusion\n## Bridging drug discovery and metabolomics and future directions","[{\"question\":\"Why is untargeted metabolomics analysis challenging?\",\"answer\":\"Large metabolomics datasets are difficult to analyze because existing approaches often depend on metabolic pathway knowledge that is limited and incompletely annotated, varying across databases.\"},{\"question\":\"How does the study use machine learning to go beyond known pathways?\",\"answer\":\"It encodes metabolite structures with molecular fingerprints and trains machine-learning models to predict experimental responses, then interprets predictions via feature-importance analysis to find chemical configurations and structurally related metabolite groups.\"},{\"question\":\"What datasets are used to validate the approach?\",\"answer\":\"The approach is tested on two datasets: one focused on Ataxia Telangiectasia and another involving endothelial cells under low oxygen, demonstrating effective prediction and pathway-consistent feature importance.\"}]","Exploring machine learning for untargeted metabolomics using molecular fingerprints | PDF",1785681584,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"exploring-machine-learning-for-untargeted-metabolomics-using-molecular-fingerprints","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/exploring-machine-learning-for-untargeted-metabolomics-using-molecular-fingerprints/118098/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05","2026-08-02",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why is untargeted metabolomics analysis challenging?","Question",{"text":76,"@type":77},"Large metabolomics datasets are difficult to analyze because existing approaches often depend on metabolic pathway knowledge that is limited and incompletely annotated, varying across databases.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the study use machine learning to go beyond known pathways?",{"text":81,"@type":77},"It encodes metabolite structures with molecular fingerprints and trains machine-learning models to predict experimental responses, then interprets predictions via feature-importance analysis to find chemical configurations and structurally related metabolite groups.",{"name":83,"@type":74,"acceptedAnswer":84},"What datasets are used to validate the approach?",{"text":85,"@type":77},"The approach is tested on two datasets: one focused on Ataxia Telangiectasia and another involving endothelial cells under low oxygen, demonstrating effective prediction and pathway-consistent feature importance.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]