[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118562-en":3,"doc-seo-118562-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118562,1374391974564,"Clementine","https://ap-avatar.wpscdn.com/avatar/14000253aa45c000a9e?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779874745381141002",8,"Research & Report","A framework to evaluate machine learning crystal stability predictions","The rapid adoption of machine learning across scientific fields requires community best practices, agreed benchmarking tasks, and robust metrics. Matbench Discovery is introduced as an evaluation framework for ML energy models, used as pre-filters on first-principles computed data in high-throughput searches for stable inorganic crystals. The work clarifies mismatches between thermodynamic stability and formation energy, and between retrospective and prospective benchmarking for materials discovery. It also provides a Python package, an adaptive leaderboard, and highlights misalignment between regression and task-relevant classification metrics, including false-positive risks near 0 eV/atom above the convex hull. Benchmarking shows universal interatomic potentials can efficiently pre-screen stable hypothetical materials.","UC Berkeley  \nUC Berkeley Previously Published Works  \nTitle  \nA framework to evaluate machine learning crystal stability predictions  \nPermalink  \n[https://escholarship.org/uc/item/7w07r8s8](https://escholarship.org/uc/item/7w07r8s8)  \nJournal  \nNature Machine Intelligence, 7(6)  \nISSN  \n2522-5839  \nAuthors  \nRiebesell, Janosh  \nGoodall, Rhys EA Benner, Philippet al.  \nPublication Date  \n2025-06-01  \nDOI  \n10.1038/s42256-025-01055-1  \nCopyright Information  \nThis work is made available under the terms of a Creative Commons Attribution License, available at [https://creativecommons.org/licenses/by/4.0/](https://creativecommons.org/licenses/by/4.0/)  \nPeer reviewed  \n[eScholarship.org](eScholarship.org) Powered by the California Digital Library  \nUniversity of California  \nA framework to evaluate machine learning crystal stability  \npredictions  \nJanosh Riebesell 1,2 Rhys E. A. Goodall 1 Philipp Benner3 Yuan Chiang2,4 Bowen Deng2,4 Gerbrand Ceder2,4 Mark Asta2,4  \nAlpha A. Lee 1 Anubhav Jain2 Kristin A. Persson2,4  \nJuly 24, 2025  \n1 Department of Physics, University of Cambridge, UK  \n2 Lawrence Berkeley National Laboratory, USA  \n3 Federal Institute of Materials Research and Testing (BAM), Germany  \n4 Department of Materials Science and Engineering, University of California-Berkeley, USA  \nAbstract  \nThe rapid adoption of machine learning in various scientific domains calls for the development of best practices and community agreed-upon benchmarking tasks and metrics. We present Matbench Discovery as an example evaluation framework for machine learning (ML) energy models, here applied as pre-filters to first-principles computed data in a high-throughput search for stable inorganic crystals. We address the disconnect between (i) thermodynamic stability and formation energy and (ii) retrospective and prospective benchmarking for materials discovery. Alongside this paper, we publish a Python package to aid with future model submissions and a growing online leaderboard with adaptive user-defined weighting of various performance metrics allowing researchers to prioritize the metrics they value most. To answer the question of which ML methodology performs best at materials discovery, our initial release includes random forests, graph neural networks (GNN), one-shot predictors, iterative Bayesian optimizers and universal interatomic potentials (UIP) . We highlight a misalignment between commonly used regression metrics and more taskrelevant classification metrics for materials discovery. Accurate regressors are susceptible to unexpectedly high false-positive rates if those accurate predictions lie close to the decision boundary at 0eV/atom above the convex hull. The benchmark results demonstrate that UIPs have advanced sufficiently to effectively and cheaply pre-screen thermodynamic stable hypothetical materials in future expansions of high-throughput materials databases.  \n1 Introduction  \nThe challenge of evaluating, benchmarking and then applying the rapid evolution of machine learning models is common across scientific domains. Specifically, the lack of agreed-upon tasks and data sets can obscure the performance of the model, making comparisons difficult. Material science is one such domain, where in the last decade, the number of ML publications and associated models have increased dramatically. Similar to other domains, such as drug discovery and protein design, the ultimate success is often associated with the discovery of a new material with specific functionality. In the combinatorial sense, material science can be viewed as an optimization problem of mixing and arranging different atoms with a merit function that captures the complex range of proper-  \n0 Correspondence to [janosh.riebesell@gmail.com](janosh.riebesell@gmail.com), kristinpers[son@berkeley.edu](son@berkeley.edu)  \nties that emerge. To date, ∼ 105 combinations have been tested experimentally [1, 2], ∼ 107 have been simulated[3– 7], and upwards of ∼ 1010 possib","cbCaiiJxzi8Vqjhx","https://ap.wps.com/l/cbCaiiJxzi8Vqjhx","pdf",4625158,1,37,"English","en",105,"# Introduction\n## Evaluation, benchmarking, and model comparison challenges\n## Materials discovery as an optimization problem\n## Computational discovery pipeline and ML advantages","[{\"question\":\"What is Matbench Discovery in this paper?\",\"answer\":\"Matbench Discovery is presented as an evaluation framework for machine learning energy models, applied as pre-filters to first-principles computed data in high-throughput searches for stable inorganic crystals.\"},{\"question\":\"What disconnect does the paper address in benchmarking materials discovery models?\",\"answer\":\"It addresses the disconnect between thermodynamic stability and formation energy, and between retrospective and prospective benchmarking for materials discovery.\"},{\"question\":\"Why can accurate regression metrics still yield poor decision outcomes?\",\"answer\":\"Accurate regressors can produce unexpectedly high false-positive rates when predictions lie close to the decision boundary at 0 eV/atom above the convex hull, where classification outcomes matter more.\"}]","A framework to evaluate machine learning crystal stability predictions | PDF",1785684219,93,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"a-framework-to-evaluate-machine-learning-crystal-stability-predictions","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/a-framework-to-evaluate-machine-learning-crystal-stability-predictions/118562/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is Matbench Discovery in this paper?","Question",{"text":75,"@type":76},"Matbench Discovery is presented as an evaluation framework for machine learning energy models, applied as pre-filters to first-principles computed data in high-throughput searches for stable inorganic crystals.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What disconnect does the paper address in benchmarking materials discovery models?",{"text":80,"@type":76},"It addresses the disconnect between thermodynamic stability and formation energy, and between retrospective and prospective benchmarking for materials discovery.",{"name":82,"@type":73,"acceptedAnswer":83},"Why can accurate regression metrics still yield poor decision outcomes?",{"text":84,"@type":76},"Accurate regressors can produce unexpectedly high false-positive rates when predictions lie close to the decision boundary at 0 eV/atom above the convex hull, where classification outcomes matter more.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]