[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123182-en":3,"doc-seo-123182-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123182,13056703019404,"Miles","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Coverage bias in small molecule machine learning - Investigating biomolecular dataset representativeness","Coverage bias limits how well small molecule machine learning models generalize beyond their training data. The study examines whether large-scale datasets sufficiently cover the space of known biomolecular structures, addressing the frequently assumed but rarely tested premise of representative training and evaluation distributions. A Maximum Common Edge Subgraph–based distance measure is introduced and paired with an efficient Integer Linear Programming and heuristic-bounds approach. Results show many widely used datasets lack uniform coverage, reducing predictive power, and two additional methods help detect divergence for future dataset design.","|  |  |  |  |\n| --- | --- | --- | --- |\n| Article [https://doi.org/10.1038/s41467-024-55462-w](https://doi.org/10.1038/s41467-024-55462-w) |  |  |  |\n| Coverage bias in small molecule machine learning |  |  |  |\n| Received: 11 September 2023\u003Cbr>Accepted: 12 December 2024 Check for updates | Fleming Kretschmer 1, Jan Seipp2, Marcus Ludwig 1,3, Gunnar W. Klau 2 & Sebastian Böcker 1 \u003Cbr>Small molecule machine learning aims to predict chemical, biochemical, or biological properties from molecular structures, with applications such as toxicity prediction, ligand binding, and pharmacokinetics. A recent trend is developing end-to-end models that avoid explicit domain knowledge. These models assume no coverage bias in training and evaluation data, meaning the data are representative of the true distribution. However, the domain of applicability is rarely considered in such models. Here, we investigate how well large-scale datasets cover the space of known biomolecular structures. For doing so, we propose a distance measure based on solving the Maximum Common Edge Subgraph (MCES) problem, which aligns well with chemical similarity. Although this method is computationally hard, we introduce an efﬁcient approach combining Integer Linear Programming and heuristic bounds. Our ﬁndings reveal that many widely-used datasets lack uniform coverage of biomolecular structures, limiting the predictive power of models trained on them. We propose two additional methods to assess whether training datasets diverge from known molecular distributions, potentially guiding future dataset creation to improve model performance. |  |  |\n| Machine learning has been successfully used in biochemistry and chemistry for decades. We consider the task of predicting chemical, biochemical or biological properties of small molecules of biological interest from their molecular structure. A recent trend is to develop end-to-end models that avoid the explicit integration of domain knowledge via inductive bias1. Noteworthy examples are generative models for novel antibiotics2 and highly toxic small molecules3, orclassiﬁers for antibiotic activity4,5, olfactory perception6 and enzymesubstrate prediction7. The MoleculeNet paper from 2018 presents 17 medium- to large-scale datasets for molecular property prediction8. These data are frequently used in machine learning to train and evaluate new models such as graph neural networks and graphormers; notably, the paper has received more than 2000 citations in 5 years.\u003Cbr>The fact that one should not use a model outside of its domain of applicability, has been well-known in the chemometrics community. The situation may be compared to spatial bias, where one uses test |  | (and training) data from a certain geographic location, but makes claims about a model’s performance for other geographic location as well9. Yet, this problem is usually ignored when training large-scale end-to-end models for predicting molecular properties. Other models10,11 are pre-trained on larger structure datasets; yet, this cannot get around the distribution bias in training and evaluation data for individual molecular properties. Recently, words of warnings have emerged that machine learning may result in a reproducibility crisis in science9,12. In particular, the datasets from MoleculeNet have been criticized13,14. Whereas it is comparatively simple to train a machine learning model that performs well in evaluations, it is much harder to derive a model that indeed contributes to solving the underlying question.\u003Cbr>The problem of generalization within a dataset has been extensively researched. For small molecules, the widely-used scaffold split ensures that evaluation is performed for scaffolds not seen in the |  |\n\n1Chair for Bioinformatics, Institute for Computer Science, Friedrich Schiller University Jena, Jena, Germany. 2Algorithmic Bioinformatics, Institute for Computer Science, Heinrich Heine University Düsseldorf, Düsseldorf, Germany. 3Currently at","cbCaifKbpEAUI4BU","https://ap.wps.com/l/cbCaifKbpEAUI4BU","pdf",3316147,1,19,"English","en",105,"# Coverage bias in small molecule machine learning\n## Background and problem motivation\n## Domain of applicability and distribution bias\n## Dataset coverage of biomolecular structure space\n## MCES-based distance and efficient computation\n## Methods to assess dataset divergence","[{\"question\":\"What is “coverage bias” in small molecule machine learning?\",\"answer\":\"It refers to the mismatch between how molecules are represented in training/evaluation datasets and the true distribution of biomolecular structures. Models then risk being applied beyond where the data coverage is representative.\"},{\"question\":\"How do the authors measure coverage of biomolecular structures?\",\"answer\":\"They propose a distance measure derived from solving the Maximum Common Edge Subgraph (MCES) problem, which aligns well with chemical similarity.\"},{\"question\":\"What do the results show about commonly used datasets?\",\"answer\":\"Many widely used datasets do not provide uniform coverage of biomolecular structures, which limits the predictive power of models trained on them. The work also provides additional ways to detect when training data diverge from known molecular distributions.\"}]","Coverage bias in small molecule machine learning - Investigating biomolecular dataset representativeness | PDF",1785815079,48,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"coverage-bias-in-small-molecule-machine-learning-investigating-biomolecular-dataset-representativeness","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/coverage-bias-in-small-molecule-machine-learning-investigating-biomolecular-dataset-representativeness/123182/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is “coverage bias” in small molecule machine learning?","Question",{"text":75,"@type":76},"It refers to the mismatch between how molecules are represented in training/evaluation datasets and the true distribution of biomolecular structures. Models then risk being applied beyond where the data coverage is representative.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do the authors measure coverage of biomolecular structures?",{"text":80,"@type":76},"They propose a distance measure derived from solving the Maximum Common Edge Subgraph (MCES) problem, which aligns well with chemical similarity.",{"name":82,"@type":73,"acceptedAnswer":83},"What do the results show about commonly used datasets?",{"text":84,"@type":76},"Many widely used datasets do not provide uniform coverage of biomolecular structures, which limits the predictive power of models trained on them. The work also provides additional ways to detect when training data diverge from known molecular distributions.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},"General","general"]