[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118499-en":3,"doc-seo-118499-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118499,962075114765,"Quinn","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","A framework to evaluate machine learning crystal stability predictions","Machine learning is rapidly adopted across scientific fields, yet the absence of community-agreed benchmarking tasks, datasets, and metrics limits meaningful comparison and best-practice development. This work introduces Matbench Discovery as an evaluation framework for machine-learning energy models, applied as pre-filters in high-throughput searches for stable inorganic crystals. It clarifies mismatches between thermodynamic stability and formation energy and addresses challenges in retrospective and prospective benchmarking for materials discovery. The release includes a Python package and a leaderboard supporting adaptive metric weighting, and it highlights how regression metrics can yield high false positives near the 0 eV/atom decision boundary.","nature machine intelligence  \nArticle [https://doi.org/10.1038/s42256-025-01055-1](https://doi.org/10.1038/s42256-025-01055-1)  \n\n| A framework to evaluate machine learning crystal stability predictions |  |  |\n| --- | --- | --- |\n| Received: 2 February 2024\u003Cbr>Accepted: 9 May 2025\u003Cbr>Published online: 23 June 2025  Check for updates | Janosh Riebesell  1,2 , Rhys E. A. Goodall  1, Philipp Benner  3, Yuan Chiang2,4, Bowen Deng  2,4, Gerbrand Ceder  2,4, Mark Asta2,4, Alpha A. Lee1, Anubhav Jain  2 & KristinA. Persson  2,4 \u003Cbr>The rapid adoption of machine learning in various scientific domains calls for the development of best practicesand community agreed-upon benchmarking tasks and metrics. We present Matbench Discovery asan example evaluation framework for machine learning energy models, here applied as pre-filters to first-principles computed data ina high-throughput search for stable inorganic crystals. We address the disconnect between (1) thermodynamic stability and formation energy and (2) retrospective and prospective benchmarking for materials discovery. Alongside this paper, we publish a Python package to aid with future model submissions anda growing online leaderboard with adaptive user-defined weighting of various performance metrics allowing researchers to prioritize the metrics they value most. To answer the question of which machine learning methodology performs best at materials discovery, our initial release includes random forests, graph neural networks, one-shot predictors, iterative Bayesian optimizersand universal interatomic potentials. We highlight amisalignment between commonly used regression metricsand more task-relevant classification metrics for materials discovery. Accurate regressors are susceptible to unexpectedly high false-positive rates if those accurate predictions lie closetothe decision boundary at 0 eV per atom above the convex hull. The benchmark results demonstrate that universal interatomic potentialshave advanced sufficiently to effectively and cheaply pre-screen thermodynamic stable hypothetical materials in future expansions of high-throughput materials databases. |  |\n| The challenge of evaluating, benchmarking and then applying the rapid evolution of machine learning (ML) models is common across scientific domains. Specifically, the lack of agreed-upon tasks and datasets can obscure the performance of the model, making comparisons difficult. Materials science is one such domain, where in the last decade, the numbers of ML publications and associated models have increased dramatically. Similar to other domains, such as drug discovery and protein design, the ultimate success is often associated |  | with the discovery of anew material with specific functionality. In the combinatorial sense, materials science can be viewed as an optimization problem of mixing and arranging different atoms with a merit function that captures the complex range of properties that emerge. To date,~105 combinations have been tested experimentally1,2, ~107 have been simulated3–7and upwards of~1010 possible quaternary materials are allowed by electronegativity and charge-balancing rules8. The space of quinternaries and higher is even less explored, leaving vast numbers |\n| 1Department of Physics, University of Cambridge, Cambridge, UK. 2Lawrence Berkeley National Laboratory, Berkeley, CA, USA. 3Federal Institute of Materials Research and Testing (BAM), Berlin, Germany. 4Department of Materials Science and Engineering, University of California, Berkeley, Berkeley, CA, USA. [e-mail: janosh.riebesell@gmail.com](e-mail: janosh.riebesell@gmail.com); [kristinpersson@berkeley.edu](kristinpersson@berkeley.edu) |  |  |\n\nNature Machine Intelligence | Volume 7 | June 2025 | 836–847 836  \nof potentially useful materials tobe discovered. The discovery of new materialsisakey driver of technological progress and lies on the path to more efficient solar cells, lighter and longer-lived batteries, and smaller and more effic","cbCairB9VoT0lfG4","https://ap.wps.com/l/cbCairB9VoT0lfG4","pdf",1413567,1,12,"English","en",105,"# Introduction\n## Need for agreed benchmarking in ML\n## Computational materials discovery challenges\n# Matbench Discovery framework\n## Energy model pre-filters for stable crystals\n## Bridging stability vs formation energy\n## Retrospective and prospective benchmarking\n# Benchmark design and evaluation metrics\n## Misalignment of regression and classification metrics\n## False positives near convex-hull boundary\n# Models, tooling, and leaderboard\n## Included ML methodologies\n## Python package and adaptive metric weighting","[{\"question\":\"What is Matbench Discovery used for?\",\"answer\":\"Matbench Discovery evaluates machine-learning energy models and applies them as pre-filters in high-throughput searches for stable inorganic crystals.\"},{\"question\":\"Why do regression metrics sometimes fail in materials discovery?\",\"answer\":\"Accurate regression can still produce unexpectedly high false-positive rates when predictions lie close to the decision boundary at 0 eV per atom above the convex hull.\"},{\"question\":\"What tools and community features accompany the framework?\",\"answer\":\"A Python package is released to support future model submissions, along with an online leaderboard that uses adaptive user-defined weighting across performance metrics.\"}]","A framework to evaluate machine learning crystal stability predictions | PDF",1785683891,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"a-framework-to-evaluate-machine-learning-crystal-stability-predictions","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/a-framework-to-evaluate-machine-learning-crystal-stability-predictions/118499/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is Matbench Discovery used for?","Question",{"text":75,"@type":76},"Matbench Discovery evaluates machine-learning energy models and applies them as pre-filters in high-throughput searches for stable inorganic crystals.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why do regression metrics sometimes fail in materials discovery?",{"text":80,"@type":76},"Accurate regression can still produce unexpectedly high false-positive rates when predictions lie close to the decision boundary at 0 eV per atom above the convex hull.",{"name":82,"@type":73,"acceptedAnswer":83},"What tools and community features accompany the framework?",{"text":84,"@type":76},"A Python package is released to support future model submissions, along with an online leaderboard that uses adaptive user-defined weighting across performance metrics.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]