[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-127041-en":3,"doc-seo-127041-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},127041,962084928904,"Asher","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","MISATO - machine learning dataset of protein–ligand complexes for structure-based drug discovery","MISATO provides a rigorously validated machine learning dataset for structure-based drug discovery by integrating quantum mechanical properties of small molecules with molecular dynamics simulations for about 20,000 experimental protein–ligand complexes. Starting from existing experimental structures, semi-empirical quantum mechanics is used to systematically refine them and a large set of explicit-water MD trajectories is included, totaling over 170 μs. The resource includes corrected biomolecule–ligand nomenclature and supports AI baseline models that improve prediction accuracy, offering an accessible entry point for next-generation drug discovery models.","nature computational science  \nResource [https://doi.org/10.1038/s43588-024-00627-2](https://doi.org/10.1038/s43588-024-00627-2)  \nMISATO: machine learning dataset of protein–ligand complexes for structure-based drug discovery  \nReceived: 30 May 2023  \nAccepted: 11 April 2024  \n\n| Published online: 10 May 2024 |\n| --- |\n|  Check for updates |\n\nTill Siebenmorgen  1,2,9, Filipe Menezes  1,2,9, Sabrina Benassou3,  \nErinc Merdivan4, Kieran Didi  5, André Santos Dias Mourão1,2, Radosław Kitel6, Pietro Liò5, Stefan Kesselheim  3, Marie Piraud4, Fabian J. Theis  4,7,8, Michael Sattler  1,2 & Grzegorz M. Popowicz  1,2   \nLarge language models have greatly enhanced our ability to understand biology and chemistry, yet robust methods for structure-based drug discovery, quantum chemistry and structural biology are still sparse.  \nPrecise biomolecule–ligand interaction datasets are urgently needed for large language models. To address this, we present MISATO, a dataset that combines quantum mechanical properties of small molecules and associated molecular dynamics simulations of~20,000 experimental protein–ligand complexes with extensive validation of experimental data. Starting from the existing experimental structures, semi-empirical quantum mechanics was used to systematically refine these structures. A large collection of molecular dynamics traces of protein–ligand complexes in explicit water is included, accumulating over 170 μs. We give examples of machine learning (ML) baseline models proving an improvement of accuracy by employing our data. An easy entry point forML experts is provided to enable the next generation of drug discovery artificial intelligence models.  \nIn recent years, artificial intelligence (AI) predictions have revolutionized many fields of science. In structural biology, AlphaFold2 (ref. 1) predicts accurate protein structures from amino-acid sequences only. Its accuracy nears state-of-the-art experimental data. The success of AlphaFold2 is made possible due toa rich database of nearly 200,000 protein structures that have been deposited and are available in the Protein DataBank (PDB)2. These structures were determined over the past decades using X-ray crystallography, nuclear magnetic resonance (NMR) or cryo-electron microscopy. Despite enormous investments,  \nthere are still few new drugs approved yearly, with development costs reaching several billion dollars3. An ongoing grand challenge is rational, structure-based drug discovery (DD). Compared with protein structure prediction, this task is substantially more difficult.  \nIn the early stages of DD, structure-based methods are popular and efficient approaches. The biomolecule provides the starting point for rational ligand search. Later, it guides optimization to optimally explore the chemical combinatorial space4 while still ensuring druglike properties. In silico methods that are in principle able to tackle  \n1Molecular Targets and Therapeutics Center, Institute of Structural Biology, Helmholtz Munich, Neuherberg, Germany. 2TUM School of Natural Sciences, Department of Bioscience, Bayerisches NMR Zentrum, Technical University of Munich, Garching, Germany. 3Jülich Supercomputing Centre, Forschungszentrum Jülich, Jülich, Germany. 4Helmholtz AI, Helmholtz Munich, Neuherberg, Germany. 5Computer Laboratory, Cambridge University, Cambridge, UK. 6Faculty of Chemistry, Jagiellonian University, Krakow, Poland. 7Computational Health Center, Institute of Computational Biology, Helmholtz Munich, Neuherberg, Germany. 8TUM School of Computation, Information and Technology, Technical University of Munich, Garching, Germany. 9These authors contributed equally: Till Siebenmorgen, Filipe Menezes. [e-mail:](e-mail: grzegorz.popowicz@helmholtz-munich.de)[ grzegorz.popowicz@helmholtz-munich.de](e-mail: grzegorz.popowicz@helmholtz-munich.de)  \nFig. 1 | MISATO combines QM data with MD-derived protein–ligand  \ndynamics. a, We provide a dataset that combines semi-empirical QM propert","cbCaidx0s5llJU23","https://ap.wps.com/l/cbCaidx0s5llJU23","pdf",2053612,1,12,"English","en",105,"# MISATO dataset overview\n## Integration of QM properties and MD simulations\n## Dataset scale and refinement protocol\n## Use of ML baselines for structure-based drug discovery","[{\"question\":\"What is MISATO designed to provide for structure-based drug discovery?\",\"answer\":\"MISATO is a curated machine learning dataset that combines semi-empirical quantum mechanical properties of small molecules with molecular dynamics simulations for experimental protein–ligand complexes, including extensive validation of experimental data.\"},{\"question\":\"How many protein–ligand complexes and how much MD simulation time does the dataset include?\",\"answer\":\"MISATO includes data for about 20,000 experimental protein–ligand complexes and provides molecular dynamics traces totaling over 170 μs.\"},{\"question\":\"How does MISATO improve structure quality before adding simulation and AI-ready features?\",\"answer\":\"Starting from existing experimental structures, it uses semi-empirical quantum mechanics to systematically refine them and includes extensive correction of common issues such as protonation and geometry, ensuring cleaner inputs for downstream modeling.\"}]","MISATO - machine learning dataset of protein–ligand complexes for structure-based drug discovery | PDF",1785936505,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"misato-machine-learning-dataset-of-proteinligand-complexes-for-structure-based-drug-discovery","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/misato-machine-learning-dataset-of-proteinligand-complexes-for-structure-based-drug-discovery/127041/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is MISATO designed to provide for structure-based drug discovery?","Question",{"text":75,"@type":76},"MISATO is a curated machine learning dataset that combines semi-empirical quantum mechanical properties of small molecules with molecular dynamics simulations for experimental protein–ligand complexes, including extensive validation of experimental data.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How many protein–ligand complexes and how much MD simulation time does the dataset include?",{"text":80,"@type":76},"MISATO includes data for about 20,000 experimental protein–ligand complexes and provides molecular dynamics traces totaling over 170 μs.",{"name":82,"@type":73,"acceptedAnswer":83},"How does MISATO improve structure quality before adding simulation and AI-ready features?",{"text":84,"@type":76},"Starting from existing experimental structures, it uses semi-empirical quantum mechanics to systematically refine them and includes extensive correction of common issues such as protonation and geometry, ensuring cleaner inputs for downstream modeling.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]