[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-126817-en":3,"doc-seo-126817-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},126817,1099523885074,"Ivy","https://ap-avatar.wpscdn.com/davatar_9964176cb1d06d4a9deccf72a44ae3dc",8,"Research & Report","MISATO - machine learning dataset of protein–ligand complexes for structure-based drug discovery","MISATO提供面向结构基础药物发现的蛋白–配体复杂物机器学习数据集，旨在支持大语言模型所需的高质量生物大分子–配体相互作用表征。数据整合了约2万组实验蛋白–配体复合物的量子力学小分子性质，并配套显式水环境下的分子动力学模拟，累计超过170 μs，同时对实验数据进行充分验证。以既有实验结构为起点，采用半经验量子力学系统性细化，并提供可用于训练与评估的机器学习基线示例，便于新一代药物发现AI建模。","nature computational science  \nResource [https://doi.org/10.1038/s43588-024-00627-2](https://doi.org/10.1038/s43588-024-00627-2)  \nMISATO: machine learning dataset of protein–ligand complexes for structure-based drug discovery  \nReceived: 30 May 2023  \nAccepted: 11 April 2024  \n\n| Published online: 10 May 2024 |\n| --- |\n|  Check for updates |\n\nTill Siebenmorgen  1,2,9, Filipe Menezes  1,2,9, Sabrina Benassou3,  \nErinc Merdivan4, Kieran Didi  5, André Santos Dias Mourão1,2, Radosław Kitel6, Pietro Liò5, Stefan Kesselheim  3, Marie Piraud4, Fabian J. Theis  4,7,8, Michael Sattler  1,2 & Grzegorz M. Popowicz  1,2   \nLarge language models have greatly enhanced our ability to understand biology and chemistry, yet robust methods for structure-based drug discovery, quantum chemistry and structural biology are still sparse.  \nPrecise biomolecule–ligand interaction datasets are urgently needed for large language models. To address this, we present MISATO, a dataset that combines quantum mechanical properties of small molecules and associated molecular dynamics simulations of~20,000 experimental protein–ligand complexes with extensive validation of experimental data. Starting from the existing experimental structures, semi-empirical quantum mechanics was used to systematically refine these structures. A large collection of molecular dynamics traces of protein–ligand complexes in explicit water is included, accumulating over 170 μs. We give examples of machine learning (ML) baseline models proving an improvement of accuracy by employing our data. An easy entry point forML experts is provided to enable the next generation of drug discovery artificial intelligence models.  \nIn recent years, artificial intelligence (AI) predictions have revolutionized many fields of science. In structural biology, AlphaFold2 (ref. 1) predicts accurate protein structures from amino-acid sequences only. Its accuracy nears state-of-the-art experimental data. The success of AlphaFold2 is made possible due toa rich database of nearly 200,000 protein structures that have been deposited and are available in the Protein DataBank (PDB)2. These structures were determined over the past decades using X-ray crystallography, nuclear magnetic resonance (NMR) or cryo-electron microscopy. Despite enormous investments,  \nthere are still few new drugs approved yearly, with development costs reaching several billion dollars3. An ongoing grand challenge is rational, structure-based drug discovery (DD). Compared with protein structure prediction, this task is substantially more difficult.  \nIn the early stages of DD, structure-based methods are popular and efficient approaches. The biomolecule provides the starting point for rational ligand search. Later, it guides optimization to optimally explore the chemical combinatorial space4 while still ensuring druglike properties. In silico methods that are in principle able to tackle  \n1Molecular Targets and Therapeutics Center, Institute of Structural Biology, Helmholtz Munich, Neuherberg, Germany. 2TUM School of Natural Sciences, Department of Bioscience, Bayerisches NMR Zentrum, Technical University of Munich, Garching, Germany. 3Jülich Supercomputing Centre, Forschungszentrum Jülich, Jülich, Germany. 4Helmholtz AI, Helmholtz Munich, Neuherberg, Germany. 5Computer Laboratory, Cambridge University, Cambridge, UK. 6Faculty of Chemistry, Jagiellonian University, Krakow, Poland. 7Computational Health Center, Institute of Computational Biology, Helmholtz Munich, Neuherberg, Germany. 8TUM School of Computation, Information and Technology, Technical University of Munich, Garching, Germany. 9These authors contributed equally: Till Siebenmorgen, Filipe Menezes. [e-mail:](e-mail: grzegorz.popowicz@helmholtz-munich.de)[ grzegorz.popowicz@helmholtz-munich.de](e-mail: grzegorz.popowicz@helmholtz-munich.de)  \nFig. 1 | MISATO combines QM data with MD-derived protein–ligand  \ndynamics. a, We provide a dataset that combines semi-empirical QM propert","cbCaii5Ny4wKmJUI","https://ap.wps.com/l/cbCaii5Ny4wKmJUI","pdf",3254345,1,14,"English","en",105,"# MISATO数据集概述\n## 数据来源与QM/MD联合构建\n## 实验验证与结构细化流程\n# 结构基础药物发现背景与挑战\n## 与AlphaFold2等方法对比\n## 传统计算方法的局限与成本\n# AI在药物发现中的机遇与现阶段不足\n## SMILES与3D数据建模差距\n# 数据集带来的ML基线示例\n## 提升精度与使用入口","[{\"question\":\"MISATO数据集整合了哪些类型的数据？\",\"answer\":\"MISATO结合了小分子的半经验量子力学性质与约2万组实验蛋白–配体复合物的分子动力学模拟数据，并包含对实验数据的扩展验证。\"},{\"question\":\"MISATO中的分子动力学模拟覆盖了什么规模与环境？\",\"answer\":\"模拟覆盖蛋白–配体复合物的全部体系，且在显式水环境中生成大量轨迹，累计时长超过170 μs。\"},{\"question\":\"MISATO如何帮助结构基础药物发现中的机器学习建模？\",\"answer\":\"数据集基于细化后的实验结构提供训练所需的3D相互作用信息，并通过机器学习基线模型展示了使用该数据集可提升准确性，从而为下一代药物发现AI提供易用入口。\"}]","MISATO - machine learning dataset of protein–ligand complexes for structure-based drug discovery | PDF",1785934966,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"misato-machine-learning-dataset-of-proteinligand-complexes-for-structure-based-drug-discovery","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/misato-machine-learning-dataset-of-proteinligand-complexes-for-structure-based-drug-discovery/126817/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"MISATO数据集整合了哪些类型的数据？","Question",{"text":75,"@type":76},"MISATO结合了小分子的半经验量子力学性质与约2万组实验蛋白–配体复合物的分子动力学模拟数据，并包含对实验数据的扩展验证。","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"MISATO中的分子动力学模拟覆盖了什么规模与环境？",{"text":80,"@type":76},"模拟覆盖蛋白–配体复合物的全部体系，且在显式水环境中生成大量轨迹，累计时长超过170 μs。",{"name":82,"@type":73,"acceptedAnswer":83},"MISATO如何帮助结构基础药物发现中的机器学习建模？",{"text":84,"@type":76},"数据集基于细化后的实验结构提供训练所需的3D相互作用信息，并通过机器学习基线模型展示了使用该数据集可提升准确性，从而为下一代药物发现AI提供易用入口。","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]