[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122129-en":3,"doc-seo-122129-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122129,1099514068035,"Ezra","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","MF-PCBA - Multifidelity High-Throughput Screening Benchmarks for Drug Discovery and Machine Learning","High-throughput screening (HTS) is widely used in drug discovery to identify drug candidates efficiently, relying on large, diverse compound libraries for high-volume activity measurements. Such data benefit computational and experimental efforts when combined with modern deep learning, but many public machine-learning-ready datasets overlook the multiple modalities present in real HTS. The multifidelity gap means noisy primary measurements are often ignored by current models. This work introduces MF-PCBA, a curated collection of 60 PubChem BioAssay datasets with primary and confirmatory modalities, enabling learning across low- and high-fidelity screens and supporting more realistic integration benchmarks.","[pubs.acs.org/jcim](pubs.acs.org/jcim)  Article   \nMF-PCBA: Multifidelity High-Throughput Screening Benchmarks for Drug Discovery and Machine Learning  \nDavid Buterez, * Jon Paul Janet, Steven J. Kiddle, and Pietro Lìo  \n Cite This: J. Chem. Inf. Model. 2023, 63, 2667−2678  \nRead Online  \nACCESS  \n Metrics & More  \n Article Recommendations  \n*sı   \nSupporting Information  \nDownloaded via UNIVERSITA DI ROMA LA SAPIENZA on November 6, 2024 at 15:50:15 (UTC) . See [https://pubs.acs.org/sharingguidelines](https://pubs.acs.org/sharingguidelines) for options on how to legitimately share published articles.  \nABSTRACT: High-throughput screening (HTS), as one of the key techniques in drug discovery, is frequently used to identify promising drug candidates in a largely automated and cost-effective way. One of the necessary conditions for successful HTS campaigns is a large and diverse compound library, enabling hundreds of thousands of activity measurements per project. Such collections of data hold great promise for computational and experimental drug discovery efforts, especially when leveraged in combination with modern deep learning techniques, and can potentially lead to improved drug activity predictions and cheaper and more effective experimental design. However, existing collections of machine-learning-ready public datasets do not exploit the multiple data modalities present in real-world HTS projects. Thus, the largest fraction of experimental measurements, corresponding to hundreds of thousands of “noisy” activity values from primary screening, are effectively ignored in the majority of machine learning models of HTS data. To address these limitations, we introduce Multifidelity PubChem BioAssay (MF-PCBA), a curated collection of 60 datasets that includes two data modalities for each dataset, corresponding to primary and confirmatory screening, an aspect that we call multifidelity. Multifidelity data accurately reflect real-world HTS conventions and present a new, challenging task for machine learning: the integration of low-and high-fidelity measurements through molecular representation learning, taking into account the orders-of-magnitude difference in size between the primary and confirmatory screens. Here we detail the steps taken to assemble MF-PCBA in terms of data acquisition from PubChemand the filtering steps required to curate the raw data. We also provide an evaluation of a recent deep-learning-based method for multifidelity integration across the introduced datasets, demonstrating the benefit of leveraging all HTS modalities, and a discussion in terms of the roughness of the molecular activity landscape. In total, MF-PCBA contains over 16.6 million unique molecule − protein interactions. The datasets can be easily assembled by using the source code available at [https://github.com/davidbuterez/](https://github.com/davidbuterez/)[ ](https://github.com/davidbuterez/)[mf-pcba.](mf-pcba.)  \n■ INTRODUCTION  \nMachine learning (ML) techniques have enabled remarkable progress in the chemical and physical sciences, particularly in terms offast and precise modeling of computationally expensive processes. Graph neural networks (GNNs), a class of geometric deep learning algorithms, have recently emerged as one of the leading ML paradigms for learning directly on the data types occurring in the life sciences. Thanks to their ability to naturally learn from non-Euclidean data structures, represented as objects (nodes) and their connections (edges), GNNs have the potential to model complex relationships and dependencies between nodes. Successful examples include but are not limited to tasks from particle physics,1 simulations of fluid dynamics and other physical systems,2 quantum chemistry,3 and drug discovery.4 However, such data-driven efforts are highly dependent on the quality, quantity, and availability of suitable data. Thus, a significant amount of research has been devoted to the development of high-quality datasets tha","cbCaijrN1y5vsSp0","https://ap.wps.com/l/cbCaijrN1y5vsSp0","pdf",3185768,1,12,"English","en",105,"# Abstract\n# Introduction\n## Machine learning for chemical sciences\n## Graph neural networks and data quality\n## Public molecular benchmarks and variability\n## Dataset diversity and benchmark examples","[{\"question\":\"What problem does MF-PCBA address in existing HTS machine-learning datasets?\",\"answer\":\"Existing public datasets typically do not exploit multiple data modalities from real HTS projects, causing noisy primary screening measurements to be largely ignored by most HTS ML models.\"},{\"question\":\"What is MF-PCBA and how is it structured?\",\"answer\":\"MF-PCBA is a curated collection of 60 PubChem BioAssay datasets, where each dataset contains two modalities corresponding to primary and confirmatory screening, forming a multifidelity setting.\"},{\"question\":\"Why is multifidelity integration important for machine learning in HTS?\",\"answer\":\"Primary and confirmatory screens differ by orders of magnitude in size, so integrating low- and high-fidelity measurements creates a challenging learning task that more accurately reflects real HTS conventions.\"}]","MF-PCBA - Multifidelity High-Throughput Screening Benchmarks for Drug Discovery and Machine Learning | PDF",1785808965,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"mf-pcba-multifidelity-high-throughput-screening-benchmarks-for-drug-discovery-and-machine-learning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/mf-pcba-multifidelity-high-throughput-screening-benchmarks-for-drug-discovery-and-machine-learning/122129/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does MF-PCBA address in existing HTS machine-learning datasets?","Question",{"text":75,"@type":76},"Existing public datasets typically do not exploit multiple data modalities from real HTS projects, causing noisy primary screening measurements to be largely ignored by most HTS ML models.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is MF-PCBA and how is it structured?",{"text":80,"@type":76},"MF-PCBA is a curated collection of 60 PubChem BioAssay datasets, where each dataset contains two modalities corresponding to primary and confirmatory screening, forming a multifidelity setting.",{"name":82,"@type":73,"acceptedAnswer":83},"Why is multifidelity integration important for machine learning in HTS?",{"text":84,"@type":76},"Primary and confirmatory screens differ by orders of magnitude in size, so integrating low- and high-fidelity measurements creates a challenging learning task that more accurately reflects real HTS conventions.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]