[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121681-en":3,"doc-seo-121681-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},121681,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1785132997149421697",8,"Research & Report","MF-PCBA - Multifidelity High-Throughput Screening Benchmarks for Drug Discovery and Machine Learning","High-throughput screening (HTS) enables automated, cost-effective identification of drug candidates, but successful campaigns require large, diverse compound libraries and extensive activity measurements. Existing public machine-learning datasets typically overlook multimodal HTS signals, especially noisy primary-screen measurements, limiting model performance. MF-PCBA is a curated collection of 60 PubChem BioAssay datasets combining primary and confirmatory modalities, enabling multifidelity integration tasks with low- and high-fidelity measurements. The benchmark supports molecular representation learning across major scale differences and includes 16.6M unique molecule–protein interactions, with evaluations of a deep-learning integration method.","[pubs.acs.org/jcim](pubs.acs.org/jcim)  Article   \nMF-PCBA: Multifidelity High-Throughput Screening Benchmarks for Drug Discovery and Machine Learning  \nDavid Buterez, * Jon Paul Janet, Steven J. Kiddle, and Pietro Lìo  \n Cite This: J. Chem. Inf. Model. 2023, 63, 2667−2678  \nRead Online  \n\n|  |  |  |  |  |  |\n| --- | --- | --- | --- | --- | --- |\n| ACCESS   | Metrics & More |  |  Article Recommendations |  | *sı Supporting Information |\n\nABSTRACT: High-throughput screening (HTS), as one of the key techniques in drug discovery, is frequently used to identify promising drug candidates in a largely automated and cost-effective way. One of the necessary conditions for successful HTS campaigns is a large and diverse compound library, enabling hundreds of thousands of activity measurements per project. Such collections of data hold great promise for computational and experimental drug discovery efforts, especially when leveraged in combination with modern deep learning techniques, and can potentially lead to improved drug activity predictions and cheaper and more effective experimental design. However, existing collections of machine-learning-ready public datasets do not exploit the multiple data modalities present in real-world HTS projects. Thus, the largest fraction of experimental measurements, corresponding to hundreds of thousands of “noisy” activity values from primary screening, are effectively ignored in the majority of machine learning models of HTS data. To address these limitations, we introduce Multifidelity PubChem BioAssay (MF-PCBA), a curated collection of 60 datasets that includes two data modalities for each dataset, corresponding to primary and confirmatory screening, an aspect that we call multifidelity. Multifidelity data accurately reflect real-world HTS conventions and present a new, challenging task for machine learning: the integration of low-and high-fidelity measurements through molecular representation learning, taking into account the orders-of-magnitude difference in size between the primary and confirmatory screens. Here we detail the steps taken to assemble MF-PCBA in terms of data acquisition from PubChemand the filtering steps required to curate the raw data. We also provide an evaluation of a recent deep-learning-based method for multifidelity integration across the introduced datasets, demonstrating the benefit of leveraging all HTS modalities, and a discussion in terms of the roughness of the molecular activity landscape. In total, MF-PCBA contains over 16.6 million unique molecule − protein interactions. The datasets can be easily assembled by using the source code available at [https://github.com/davidbuterez/](https://github.com/davidbuterez/)[ ](https://github.com/davidbuterez/)[mf-pcba.](mf-pcba.)  \n■ INTRODUCTION  \nMachine learning (ML) techniques have enabled remarkable progress in the chemical and physical sciences, particularly in terms offast and precise modeling of computationally expensive processes. Graph neural networks (GNNs), a class of geometric deep learning algorithms, have recently emerged as one of the leading ML paradigms for learning directly on the data types occurring in the life sciences. Thanks to their ability to naturally learn from non-Euclidean data structures, represented as objects (nodes) and their connections (edges), GNNs have the potential to model complex relationships and dependencies between nodes. Successful examples include but are not limited to tasks from particle physics,1 simulations of fluid dynamics and other physical systems,2 quantum chemistry,3 and drug discovery.4 However, such data-driven efforts are highly dependent on the quality, quantity, and availability of suitable data. Thus, a significant amount of research has been devoted to the development of high-quality datasets that support the rapidly advancing field of graph representation learning. In  \nparticular, challenging computational chemistry tasks such as quantum property prediction, ","cbCaiu8oDolRuxHW","https://ap.wps.com/l/cbCaiu8oDolRuxHW","pdf",3242753,1,12,"English","en",105,"# Abstract\n## Introduction\n## Data and Benchmark Construction\n## Evaluation of Multifidelity Integration","[{\"question\":\"What problem does MF-PCBA address in machine learning for HTS?\",\"answer\":\"MF-PCBA addresses the fact that many public machine-learning-ready HTS datasets ignore large portions of real-world multimodal data, especially noisy primary screening measurements, which are crucial for more realistic modeling.\"},{\"question\":\"What is multifidelity in the context of MF-PCBA?\",\"answer\":\"Multifidelity refers to using two data modalities per assay: primary screening and confirmatory screening. MF-PCBA frames HTS integration as combining low- and high-fidelity measurements.\"},{\"question\":\"How many datasets and molecule–protein interactions are included in MF-PCBA?\",\"answer\":\"MF-PCBA contains 60 curated datasets and over 16.6 million unique molecule–protein interactions.\"}]","MF-PCBA - Multifidelity High-Throughput Screening Benchmarks for Drug Discovery and Machine Learning | PDF",1785806183,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"mf-pcba-multifidelity-high-throughput-screening-benchmarks-for-drug-discovery-and-machine-learning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/mf-pcba-multifidelity-high-throughput-screening-benchmarks-for-drug-discovery-and-machine-learning/121681/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does MF-PCBA address in machine learning for HTS?","Question",{"text":75,"@type":76},"MF-PCBA addresses the fact that many public machine-learning-ready HTS datasets ignore large portions of real-world multimodal data, especially noisy primary screening measurements, which are crucial for more realistic modeling.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is multifidelity in the context of MF-PCBA?",{"text":80,"@type":76},"Multifidelity refers to using two data modalities per assay: primary screening and confirmatory screening. MF-PCBA frames HTS integration as combining low- and high-fidelity measurements.",{"name":82,"@type":73,"acceptedAnswer":83},"How many datasets and molecule–protein interactions are included in MF-PCBA?",{"text":84,"@type":76},"MF-PCBA contains 60 curated datasets and over 16.6 million unique molecule–protein interactions.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]