[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123472-en":3,"doc-seo-123472-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123472,687197207919,"Theodora","https://ap-avatar.wpscdn.com/avatar/a000253d6f5f7c60be?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779446848396160552",8,"Research & Report","Nutmeg and SPICE: Models and Data for Biomolecular Machine Learning","The content presents version 2 of the SPICE dataset, a quantum chemistry resource built to train machine learning potentials. The update expands chemical-space coverage with extensive new sampling, adds richer non-covalent interaction data, and increases overall dataset size. Models termed Nutmeg are trained using a TensorNet-based architecture that injects precomputed partial charges to improve charged and polar predictions, plus a repulsive term for stability. Evaluations show accurate reproduction of energy differences and stable molecular dynamics trajectories suitable for routine small-molecule simulations.","Nutmeg and SPICE: Models and Data for Biomolecular Machine Learning  \nPeter Eastman1 , Benjamin P. Pritchard2 , John D. Chodera3 , Thomas E. Markland 1  \n1 Department of Chemistry, Stanford University, Stanford, CA 94305, USA 2 Molecular Sciences Software Institute, Virginia Polytechnic Institute and State University,  \nBlacksburg, VA 24060, USA  \n3Computational and Systems Biology Program, Sloan Kettering Institute, Memorial Sloan  \nKettering Cancer Center, New York, NY 10065, USA corresponding author: Peter Eastman ([peastman@stanford.edu](peastman@stanford.edu))  \nAbstract  \nWe describe version 2 of the SPICE dataset, a collection of quantum chemistry calculations for training machine learning potentials. It expands on the original dataset by adding much more sampling of chemical space and more data on non-covalent interactions. We train a set of potential energy functions called Nutmeg on it. They are based on the TensorNet architecture. They use a novel mechanism to improve performance on charged and polar molecules, injecting precomputed partial charges into the model to provide a reference for the large scale charge distribution. Evaluation of the new models shows they do an excellent job of reproducing energy differences between conformations, even on highly charged molecules or ones that are significantly larger than the molecules in the training set. They also producestable molecular dynamics trajectories, and are fast enough to be useful for routine simulation of small molecules.  \nIntroduction  \nMachine learning potentials are a popular tool for molecular simulation.1 A large model, typically a neural network, is trained on a large body of forces and/or energies computed with a high level quantum chemistry method. The model learns to compute forces and energies for new conformations and, in some cases, new molecules. They represent a middle ground between conventional force fields, which are faster but have limited accuracy and transferability, and quantum chemistry methods, which are accurate but very slow.  \nThe two essential ingredients that make up a machine learning potential are the model  \narchitecture and the data it is trained on. Much work has recently been done to develop both of  \nthese ingredients. Many flexible, general purpose architectures have been designed that can be applied to broad areas of chemical and conformational space.2–5  \nMany quantum chemistry datasets have also been introduced, but their generality tends to be much lower. A model is unlikely to work for situations far outside the domain of the data it was trained on. Given that chemical and conformational space are nearly infinite, any dataset will necessarily be limited in the applications for which it can be used. For example, some of the most popular datasets contain only a very small number of chemical elements.6–8 Some contain only energy minimized conformations,6,9,10 which limits their usefulness for training potential functions that can be applied to molecular dynamics. Some contain only neutral molecules,7,9 limiting their usefulness for learning to simulate charged molecules. There is thus an ongoing need for more datasets that can be used for other applications.  \nSPICE is a quantum chemistry dataset designed for training machine learning potentials.11 Its focus is particularly on modeling drug-like small molecules interacting with proteins. The original version contains forces and energies for 1.1 million conformations of molecules covering a wide range of chemical space, including drug molecules, dipeptides, and solvated amino acids. There are 15 elements in total, charged and uncharged molecules, low and high energy conformations, and a wide variety of covalent and non-covalent interactions.  \nIn this article we describe an updated version of the SPICE dataset. It greatly expands the coverage of chemical space with over 20,000 new molecules. It also improves sampling of noncovalent interactions and adds two more elements","cbCainY6f1RbG2Lj","https://ap.wps.com/l/cbCainY6f1RbG2Lj","pdf",687404,1,26,"English","en",105,"# Abstract\n# Introduction\n## Machine learning potentials and data limitations\n## SPICE dataset and drug-like small-molecule focus\n# Methods\n## SPICE 2 dataset design and updates\n## Data subsets and generation approach\n## Nutmeg model architecture and training strategy","[{\"question\":\"What is SPICE dataset version 2 used for?\",\"answer\":\"SPICE v2 is a quantum chemistry dataset designed for training machine learning potentials, providing forces and energies for molecular conformations to learn accurate potential energy functions.\"},{\"question\":\"How does Nutmeg improve performance for charged and polar molecules?\",\"answer\":\"Nutmeg uses a mechanism that injects static, precomputed atomic partial charges into the model as inputs, giving a reference for large-scale charge distribution and improving accuracy for charged and polar systems.\"},{\"question\":\"What kinds of results does the updated system achieve in evaluation?\",\"answer\":\"The evaluation demonstrates accurate reproduction of energy differences between conformations and produces stable molecular dynamics trajectories, with runtime performance suitable for routine simulation of small molecules.\"}]","Nutmeg and SPICE: Models and Data for Biomolecular Machine Learning | PDF",1785816712,66,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"nutmeg-and-spice-models-and-data-for-biomolecular-machine-learning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/nutmeg-and-spice-models-and-data-for-biomolecular-machine-learning/123472/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is SPICE dataset version 2 used for?","Question",{"text":75,"@type":76},"SPICE v2 is a quantum chemistry dataset designed for training machine learning potentials, providing forces and energies for molecular conformations to learn accurate potential energy functions.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does Nutmeg improve performance for charged and polar molecules?",{"text":80,"@type":76},"Nutmeg uses a mechanism that injects static, precomputed atomic partial charges into the model as inputs, giving a reference for large-scale charge distribution and improving accuracy for charged and polar systems.",{"name":82,"@type":73,"acceptedAnswer":83},"What kinds of results does the updated system achieve in evaluation?",{"text":84,"@type":76},"The evaluation demonstrates accurate reproduction of energy differences between conformations and produces stable molecular dynamics trajectories, with runtime performance suitable for routine simulation of small molecules.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]