[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118560-en":3,"doc-seo-118560-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118560,1374391974564,"Clementine","https://ap-avatar.wpscdn.com/avatar/14000253aa45c000a9e?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779874745381141002",8,"Research & Report","A framework to evaluate machine learning crystal stability predictions - Matbench Discovery","Machine learning adoption in scientific discovery requires best practices, community benchmarking tasks, and comparable metrics. Matbench Discovery is presented as an evaluation framework for machine-learning energy models, used as pre-filters for first-principles high-throughput searches of stable inorganic crystals. The work addresses mismatches between thermodynamic stability and formation energy and between retrospective and prospective benchmarking. It also provides a Python package and a growing leaderboard with adaptive metric weighting to align evaluation with materials-discovery decision boundaries.","nature machine intelligence  \nArticle [https://doi.org/10.1038/s42256-025-01055-1](https://doi.org/10.1038/s42256-025-01055-1)  \n\n| A framework to evaluate machine learning crystal stability predictions |  |  |\n| --- | --- | --- |\n| Received: 2 February 2024\u003Cbr>Accepted: 9 May 2025\u003Cbr>Published online: 23 June 2025  Check for updates | Janosh Riebesell  1,2 , Rhys E. A. Goodall  1, Philipp Benner  3, Yuan Chiang2,4, Bowen Deng  2,4, Gerbrand Ceder  2,4, Mark Asta2,4, Alpha A. Lee1, Anubhav Jain  2 & KristinA. Persson  2,4 \u003Cbr>The rapid adoption of machine learning in various scientific domains calls for the development of best practicesand community agreed-upon benchmarking tasks and metrics. We present Matbench Discovery asan example evaluation framework for machine learning energy models, here applied as pre-filters to first-principles computed data ina high-throughput search for stable inorganic crystals. We address the disconnect between (1) thermodynamic stability and formation energy and (2) retrospective and prospective benchmarking for materials discovery. Alongside this paper, we publish a Python package to aid with future model submissions anda growing online leaderboard with adaptive user-defined weighting of various performance metrics allowing researchers to prioritize the metrics they value most. To answer the question of which machine learning methodology performs best at materials discovery, our initial release includes random forests, graph neural networks, one-shot predictors, iterative Bayesian optimizersand universal interatomic potentials. We highlight amisalignment between commonly used regression metricsand more task-relevant classification metrics for materials discovery. Accurate regressors are susceptible to unexpectedly high false-positive rates if those accurate predictions lie closetothe decision boundary at 0 eV per atom above the convex hull. The benchmark results demonstrate that universal interatomic potentialshave advanced sufficiently to effectively and cheaply pre-screen thermodynamic stable hypothetical materials in future expansions of high-throughput materials databases. |  |\n| The challenge of evaluating, benchmarking and then applying the rapid evolution of machine learning (ML) models is common across scientific domains. Specifically, the lack of agreed-upon tasks and datasets can obscure the performance of the model, making comparisons difficult. Materials science is one such domain, where in the last decade, the numbers of ML publications and associated models have increased dramatically. Similar to other domains, such as drug discovery and protein design, the ultimate success is often associated |  | with the discovery of anew material with specific functionality. In the combinatorial sense, materials science can be viewed as an optimization problem of mixing and arranging different atoms with a merit function that captures the complex range of properties that emerge. To date,~105 combinations have been tested experimentally1,2, ~107 have been simulated3–7and upwards of~1010 possible quaternary materials are allowed by electronegativity and charge-balancing rules8. The space of quinternaries and higher is even less explored, leaving vast numbers |\n| 1Department of Physics, University of Cambridge, Cambridge, UK. 2Lawrence Berkeley National Laboratory, Berkeley, CA, USA. 3Federal Institute of Materials Research and Testing (BAM), Berlin, Germany. 4Department of Materials Science and Engineering, University of California, Berkeley, Berkeley, CA, USA. [e-mail: janosh.riebesell@gmail.com](e-mail: janosh.riebesell@gmail.com); [kristinpersson@berkeley.edu](kristinpersson@berkeley.edu) |  |  |\n\nNature Machine Intelligence | Volume 7 | June 2025 | 836–847 836  \nof potentially useful materials tobe discovered. The discovery of new materialsisakey driver of technological progress and lies on the path to more efficient solar cells, lighter and longer-lived batteries, and smaller and more effic","cbCainsQ6b4FlmWS","https://ap.wps.com/l/cbCainsQ6b4FlmWS","pdf",1378742,1,12,"English","en",105,"# Introduction\n## Need for agreed-upon tasks, datasets, and metrics\n## Materials discovery as an optimization problem\n# Framework: Matbench Discovery\n## Evaluating machine-learning energy models for stable crystals\n## Handling gaps between stability and formation energy\n# Benchmarking design and tooling\n## Python package and adaptive leaderboard weighting\n## Comparing model families and task-relevant metrics\n# Key findings\n## Regression accuracy and false-positive risks near decision boundaries\n## Universal interatomic potentials for efficient pre-screening","[{\"question\":\"What is Matbench Discovery used for in this work?\",\"answer\":\"Matbench Discovery evaluates machine-learning energy models as pre-filters to accelerate first-principles high-throughput searches for stable inorganic crystals.\"},{\"question\":\"Why do the authors emphasize a disconnect between metrics and stability?\",\"answer\":\"They address mismatches between thermodynamic stability and formation energy and between regression-style benchmarking and more task-relevant classification metrics for materials discovery.\"},{\"question\":\"What tools and resources are released alongside the paper?\",\"answer\":\"A Python package is published to support future model submissions, along with a growing online leaderboard that uses adaptive user-defined weighting of performance metrics.\"}]","A framework to evaluate machine learning crystal stability predictions - Matbench Discovery | PDF",1785684204,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"a-framework-to-evaluate-machine-learning-crystal-stability-predictions-matbench-discovery","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/a-framework-to-evaluate-machine-learning-crystal-stability-predictions-matbench-discovery/118560/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05","2026-08-02",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What is Matbench Discovery used for in this work?","Question",{"text":76,"@type":77},"Matbench Discovery evaluates machine-learning energy models as pre-filters to accelerate first-principles high-throughput searches for stable inorganic crystals.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"Why do the authors emphasize a disconnect between metrics and stability?",{"text":81,"@type":77},"They address mismatches between thermodynamic stability and formation energy and between regression-style benchmarking and more task-relevant classification metrics for materials discovery.",{"name":83,"@type":74,"acceptedAnswer":84},"What tools and resources are released alongside the paper?",{"text":85,"@type":77},"A Python package is published to support future model submissions, along with a growing online leaderboard that uses adaptive user-defined weighting of performance metrics.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":122},"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]