[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-124484-en":3,"doc-seo-124484-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},124484,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","Massive Atomic Diversity - a compact universal dataset for atomistic machine learning","Machine-learning potentials for atomistic simulations increasingly rely on large databases of atomic structures and properties computed with electronic-structure methods. The MAD dataset is introduced to train models for arbitrary structures by emphasizing “massive atomic diversity” through aggressive perturbations from stable starting sets and highly consistent calculation settings. Despite fewer than 100,000 entries, MAD enables universal interatomic potentials comparable to models trained on far larger datasets and provides low-dimensional structural latent descriptors for general-purpose materials cartography.","[www. nature.com/scientificdata](www. nature.com/scientificdata)  \noPeN  \nData DeSCRiPtoR  \nMassive atomic Diversity: a compact universal dataset for atomistic machine learning  \narslan Mazitov1 ✉, Sofiia Chorna1, Guillaume Fraux1, Marnik Bercx2, Giovanni Pizzi2, Sandip De3 & Michele Ceriotti1 ✉  \nThe development of machine-learning models for atomic-scale simulations has greatly benefited from the large databases of materials and molecular properties, computed using electronic-structure calculations. Recently, these databases enabled the training of“universal” models that aim to  \nmake accurate predictions for arbitrary atomic geometries and compositions. However, many of these databases were originally designed for materials discovery, focusing primarily on equilibrium structures. Here, we introduce a dataset designed to train machine-learning models to make reasonable predictions for arbitrary structures. Starting with relatively small sets of stable structures, we built the dataset aiming to achieve “massive atomic diversity”(MaD) by aggressively modifying these structures and utilizing highly consistent electronic-structure settings for property calculations. Despite containing fewer than 100,000 entries, the MAD dataset has already enabled the training of universal interatomic potentials that rival those trained on datasets containing two to three orders of magnitude more data. We detail the design philosophy of the dataset and introduce low-dimensional structural latent space descriptors that can be used as a general-purpose materials cartography tool.  \nBackground & Summary  \nThe introduction of large-scale, open-access materials databases has significantly accelerated computational materials science and discovery1. They offer vast repositories of atomic structures and computed or experimentally measured properties of organic and inorganic compounds, facilitating high-throughput screening for many materials-discovery applications. Among these, the databases of electronic structure calculations serve as a particularly important source of data for atomistic modeling, providing a robust and consistent way of exploring structure-property relations for a wide range of materials, including those that have never been experimentally realized2–11. Despite these advancements, existing datasets primarily focus on structures at or near the local minima and saddle points of the potential energy surface (PES), limiting their applicability in atomistic simulations that often require exploration of mid-and high-energy configurations. This is particularly important for interatomic potentials—approximations ofthe PES—which require accurate descriptions of both low-and high-energy states to ensure robustness across a wide range of thermodynamic conditions. Another source of error stems from the presence of inconsistencies in computational settings between different datasets and between different structures within the dataset. For example, some compositions may be treated with different electronic-structure details to tackle known shortcomings of density-functional theory, which, however, means that different portions of chemical space are associated with different PES. Furthermore, most of the existing datasets are focused on either organic or inorganic materials – which is well motivated by the fact that these classes of materials often require different electronic-structure details and cover different energy scales, but restricts the development of universal interatomic potentials capable of handling hybrid systems of various nature and chemical compositions.  \nTo address these challenges, we introduce the Massive Atomic Diversity (MAD) dataset, designed to encompass a broad spectrum of atomic configurations, including both organic and inorganic systems, while being restricted to a small number of structures – which facilitates property estimation with converged settings, and  \n1 Laboratory of Computational Science and Modeling, Instit","cbCaibckXarglUBW","https://ap.wps.com/l/cbCaibckXarglUBW","pdf",2441566,1,12,"English","en",105,"# Background & Summary\n## Motivation from existing materials databases\n## Challenges: energy-range coverage and setting inconsistencies\n## Dataset design goals and construction\n# MAD dataset structure and components\n## Structural subsets and statistics (MC3D, MC2D, SHIFTML)\n## Element representation and diagnostics\n# Model-ready representations\n## Low-dimensional structural latent space descriptors","[{\"question\":\"What problem does the MAD dataset address in atomistic machine-learning training?\",\"answer\":\"Existing datasets often emphasize near-minimum and saddle-point structures, limiting usefulness for mid- and high-energy configurations required in simulations and for robust potential-energy-surface coverage. MAD targets reasonable predictions for arbitrary structures instead.\"},{\"question\":\"How does MAD achieve “massive atomic diversity” with relatively small dataset size?\",\"answer\":\"MAD starts from relatively small sets of stable structures, then aggressively modifies them while using highly consistent calculation settings for property evaluations. This increases structural variety without requiring orders of magnitude more entries.\"},{\"question\":\"What capabilities does MAD provide beyond training data?\",\"answer\":\"MAD introduces low-dimensional structural latent space descriptors intended as general-purpose tools for materials cartography, enabling compact representations suitable for broader exploration of materials structure–property relations.\"}]","Massive Atomic Diversity - a compact universal dataset for atomistic machine learning | PDF",1785822709,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"massive-atomic-diversity-a-compact-universal-dataset-for-atomistic-machine-learning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/massive-atomic-diversity-a-compact-universal-dataset-for-atomistic-machine-learning/124484/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the MAD dataset address in atomistic machine-learning training?","Question",{"text":75,"@type":76},"Existing datasets often emphasize near-minimum and saddle-point structures, limiting usefulness for mid- and high-energy configurations required in simulations and for robust potential-energy-surface coverage. MAD targets reasonable predictions for arbitrary structures instead.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does MAD achieve “massive atomic diversity” with relatively small dataset size?",{"text":80,"@type":76},"MAD starts from relatively small sets of stable structures, then aggressively modifies them while using highly consistent calculation settings for property evaluations. This increases structural variety without requiring orders of magnitude more entries.",{"name":82,"@type":73,"acceptedAnswer":83},"What capabilities does MAD provide beyond training data?",{"text":84,"@type":76},"MAD introduces low-dimensional structural latent space descriptors intended as general-purpose tools for materials cartography, enabling compact representations suitable for broader exploration of materials structure–property relations.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]