[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119600-en":3,"doc-seo-119600-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119600,687197207919,"Theodora","https://ap-avatar.wpscdn.com/avatar/a000253d6f5f7c60be?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779446848396160552",8,"Research & Report","Machine Learning-Assisted False Positive Detection in Metabolite Identification Workflows - Abstract and Introduction","Metabolite identification is a pivotal step in drug discovery and development, but liquid chromatography–mass spectrometry complexity often produces numerous false positives that hinder the recognition of true metabolites. The study presents a machine-learning approach that leverages expert knowledge to build discriminative features for metabolite-related chromatographic peaks using mass spectra, chromatographic signals, and kinetic profiles. Gradient boosting decision tree classifiers are validated on public and proprietary “real-world” datasets, reducing false positive identifications and improving workflow efficiency and accuracy.","This article is licensed under CC-BY-NC-ND 4.0  \n[pubs.acs.org/ac](pubs.acs.org/ac)  Article   \nMachine Learning-Assisted False Positive Detection in Metabolite Identification Workflows  \nRamon Ad̀alia, *,⊥ Paula Cifuentes,⊥ Joyce Liu, Lionel Cheruzel, Gemma Sanjuan, Tom̀as Margalef, and Ismael Zamora  \n Cite This: [https://doi.org/10.1021/acs.analchem.5c02745](https://doi.org/10.1021/acs.analchem.5c02745)  \nRead Online  \nDownloaded via UNIV AUTONOMA DE BARCELONA on December 1 1, 2025 at 07:02:30 (UTC) . See [https://pubs.acs.org/sharingguidelines](https://pubs.acs.org/sharingguidelines) for options on how to legitimately share published articles.  \nACCESS  \n Metrics & More  \n Article Recommendations  \n*sı   \nSupporting Information  \nABSTRACT: Metabolite identification is a pivotal step in drug discovery and development, enabling the comprehensive analysis of drug-derived compounds within biological systems. However, the complexity of liquid chromatography−mass spectrometry data often results in numerous false positives, complicating the identification of true metabolites. This study introduces a machine-learning-based approach to improve the accuracy of false positive detection in metabolite identification workflows. By incorporating expert knowledge, we develop a feature set for metabolite-related chromatographic peaks that characterizes true and false positives with high accuracy, integrating data from mass spectra, chromatographic signals, and kinetic profiles. We validate this method via gradient boosting decision tree classifiers on both  \npublicly available and proprietary “real-world” data sets, including small molecules and new modalities. Our findings demonstrate that machine learning-assisted techniques significantly reduce false positive identifications, thereby increasing the efficiency and accuracy of metabolite identification processes.  \n■ INTRODUCTION  \nMetabolite identification (MetID) is a critical part of drug discovery and development, providing insights into metabolic liabilities and pathways of drug candidates and aiding the identification of lead compounds for safe and effective medicines. Accurate metabolite profiling is essential not only for understanding metabolic challenges and pharmacokinetic and pharmacodynamic properties but also for meeting regulatory standards.  \nLiquid chromatography−mass spectrometry (LC-MS) is the predominant analytical technique for MetID in the pharmaceutical industry, valued for its speed, stability, sensitivity, and automation potential. It detects a wide range of metabolites in complex samples, producing large data sets visualized as chromatogram peaks. However, not all peaks represent true metabolites; false positives can arise from contamination, noise, processing errors, or even variations in LC-MS setups such as chromatography conditions, ionization methods, or mass analyzers.  \nAutomatic software tools efficiently identify major metabolite peaks based on chromatographic and spectral features, 1 typically when signal intensity is sufficient. Lower-abundance metabolites may be discarded if quality thresholds are applied, risking missed detections, since MS signal intensity does not always correlate with concentration. Consequently, experts often configure software to reveal broad candidate peaks and  \nmanually approve or reject them based on criteria including cross-peak parameters, fragments shared with parent compounds, or trends across multiple samples (e.g., incubation time series).  \nManual review is time-consuming and prone to error, introducing variability and limiting scalability as the LC-MS data volume grows with high-throughput screening.  \nTo address these limitations, we developed a machine learning framework for automatic detection and reduction of false positives in MetID. Combining advanced classification models with domain-specific feature engineering, our approach improves the reliability of software-generated annotations, while reducing manual rev","cbCaiePVIgFua0GB","https://ap.wps.com/l/cbCaiePVIgFua0GB","pdf",2381247,1,9,"English","en",105,"# Abstract\n# Introduction\n## Metabolite identification in drug discovery\n## Why LC-MS generates false positives\n## Limits of automatic and manual MetID review\n## Machine learning framework for false positive reduction\n# Related Work","[{\"question\":\"Why do false positives occur in metabolite identification workflows?\",\"answer\":\"False positives arise from contamination, noise, processing errors, or variations in LC-MS setups such as chromatography conditions, ionization methods, and mass analyzers. These factors can create peaks that do not correspond to true metabolites.\"},{\"question\":\"How does the proposed machine-learning approach improve false positive detection?\",\"answer\":\"The method incorporates expert knowledge to engineer a feature set describing true versus false metabolite-related chromatographic peaks. It combines mass spectra, chromatographic signals, and kinetic profiles to train classifiers that prioritize likely candidates.\"},{\"question\":\"Does the machine-learning model replace expert labeling in this workflow?\",\"answer\":\"No. Training labels come from manual expert review, and the models do not autonomously generate labels. The system is designed to assist experts by accelerating evaluation and reducing manual effort.\"}]","Machine Learning-Assisted False Positive Detection in Metabolite Identification Workflows - Abstract and Introduction | PDF",1785725219,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"machine-learning-assisted-false-positive-detection-in-metabolite-identification-workflows-abstract-and-introduction","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/machine-learning-assisted-false-positive-detection-in-metabolite-identification-workflows-abstract-and-introduction/119600/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do false positives occur in metabolite identification workflows?","Question",{"text":75,"@type":76},"False positives arise from contamination, noise, processing errors, or variations in LC-MS setups such as chromatography conditions, ionization methods, and mass analyzers. These factors can create peaks that do not correspond to true metabolites.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed machine-learning approach improve false positive detection?",{"text":80,"@type":76},"The method incorporates expert knowledge to engineer a feature set describing true versus false metabolite-related chromatographic peaks. It combines mass spectra, chromatographic signals, and kinetic profiles to train classifiers that prioritize likely candidates.",{"name":82,"@type":73,"acceptedAnswer":83},"Does the machine-learning model replace expert labeling in this workflow?",{"text":84,"@type":76},"No. Training labels come from manual expert review, and the models do not autonomously generate labels. The system is designed to assist experts by accelerating evaluation and reducing manual effort.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]