[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86252-en":3,"doc-seo-86252-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86252,687197207919,"Theodora","https://ap-avatar.wpscdn.com/avatar/a000253d6f5f7c60be?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779446848396160552",8,"Research & Report","AutoMatBench An Automatic Optimization Toolkit for the Acceleration of Material Properties Prediction Benchmarking","Material property prediction (MPP) infers key physical, chemical, or functional properties from chemical composition and structure, speeding discovery and optimization of novel materials. MatBench is a widely used benchmark defining more than ten representative MPP problems and evaluating in-distribution (ID) performance, but it fails to capture out-of-distribution (OOD) behavior. AutoMatBench combines MatBench benchmarking pipelines with OOD evaluation to cover extensive configuration spaces, revealing large performance discrepancies and reducing cost via Bayesian optimization within twelve optimization steps.","AutoMatBench: An Automatic Optimization Toolkit for the Acceleration of Material Properties Prediction Benchmarking Hongxiao Lia,b , Wanling Gaoa,b  \na Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China b University of Chinese Academy of Sciences, Beijing, China  \narXiv :2607 . 1 1526v 1 [ cs .LG] 13 Jul 2026  \nARTICLE INFO  \nKeywords:  \nArtificial intelligence for science  \nArtificial intelligence for materials Material properties prediction Benchmark  \nEvaluation  \nAutomatic optimization  \nAB STRACT  \nMaterial property prediction (MPP) infers key properties from chemical composition and structure, accelerating the discovery and optimization of novel materials. In the realm of MPP, MatBench is a widely accepted benchmarking tool that defines over ten significant problems and provides the paradigm of performance evaluation for AI prediction models. Even though MatBench works wellin benchmarking the performances of prediction models on in-distribution (ID) tasks and datasets, it lacks the ability to reflect their performances on out-of-distribution (OOD) material data, resulting failure in new material discovery. By combining the pipelines of MatBench and the existing researches on OOD performance evaluation, this study enables a huge space of benchmarking configurations, comprehensively reflecting the performances, abilities, and disadvantages of various AI prediction models. This work reports that the discrepancy of performances at different configuration values is huge and can be illustrated with prior knowledge and novel insights, therefore consideration of causal effect of configurations on performance results is necessary. In case of the impossibility of enumerative benchmarking at every configuration, this work further proposes AutoMatBench, an automatic toolkit with Bayesian optimization. Experiments with AutoMatBench reports that, within twelve steps of optimization, the similar results with MatBench and former OOD research can be accessed while more than half of the cost are saved. Besides, this tool also yields more essential findings on MPP benchmarking, positively contributing to the cost and efficiency of new material discovery.  \n1. Introduction  \nMaterial property prediction (MPP), as a signficant subfield of Artificial Intelligence for Science (AI4S), aims at predicting or inferring the physical, chemical, or functional properties of materials directly from their composition and structure data. By learning relationships of compositionalor structural information from experimental and computational data, MPP models enable high-efficiency screening of demanded materials, guide target-oriented synthesis, and reduces reliance on high-cost experiments or ab initio simulations.  \nMatBench [1] emerges as a widely accepted benchmarking suite that defines over ten diverse representative prediction tasks for both classification and regression problems, curated from solid datasets. MatBench provides a consistent training-test splitting and cross-fold validation, establishing a generalized and comparable paradigm for evaluating the prediction accuracies of AI models across tasks and datasets. It has become the most adopted de facto reference for developing and validating new material informatics methods.  \nUnfortunately, benchmarking results with MatBench only represent the performances on in-distribution (ID) tasks and datasets, while the performances on out-of-distribution (OOD) data are totally different. Research from Omee et al. [2] and Fung et al. [3] also support such conclusion. Experiments of this work also report similar outcomes.  \nThis study aims at establishing an evaluation methodology the abilities of AI prediction models that works both on ID and OOD data, referring the pipelines of MatBench [1] and the existing OOD performance evaluation. There are two main challenges. On the one hand, the existing OOD  \nperformance evaluation [4, 5, 6, 2] are neither completely undergone under t","cbCaicTY8Y5LotCw","https://ap.wps.com/l/cbCaicTY8Y5LotCw","pdf",2901664,5,1,21,"English","en",105,"# Introduction\n## Material property prediction and MatBench\n## Limitations of in-distribution benchmarking\n## Goal: unified ID/OOD evaluation methodology\n## Tasks, models, and configuration factors","[{\"question\":\"What limitation of MatBench motivates this work?\",\"answer\":\"MatBench benchmarks mainly in-distribution (ID) tasks and datasets, while out-of-distribution (OOD) material data yields different performance, so it cannot reflect real generalization for new material discovery.\"},{\"question\":\"How does AutoMatBench improve benchmarking beyond MatBench?\",\"answer\":\"AutoMatBench integrates MatBench’s pipeline with existing OOD performance evaluation ideas, enabling a much larger space of benchmarking configurations and more comprehensive reflection of model strengths and weaknesses.\"},{\"question\":\"Why is configuration sensitivity important in MPP benchmarking?\",\"answer\":\"Performance for the same model and task can vary greatly across configuration values, so relying on a single or a few configurations can introduce bias and misrepresent the model’s true ability.\"}]",1784209832,53,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"automatbench-an-automatic-optimization-toolkit-for-the-acceleration-of-material-properties-prediction-benchmarking","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/automatbench-an-automatic-optimization-toolkit-for-the-acceleration-of-material-properties-prediction-benchmarking/86252/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What limitation of MatBench motivates this work?","Question",{"text":76,"@type":77},"MatBench benchmarks mainly in-distribution (ID) tasks and datasets, while out-of-distribution (OOD) material data yields different performance, so it cannot reflect real generalization for new material discovery.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does AutoMatBench improve benchmarking beyond MatBench?",{"text":81,"@type":77},"AutoMatBench integrates MatBench’s pipeline with existing OOD performance evaluation ideas, enabling a much larger space of benchmarking configurations and more comprehensive reflection of model strengths and weaknesses.",{"name":83,"@type":74,"acceptedAnswer":84},"Why is configuration sensitivity important in MPP benchmarking?",{"text":85,"@type":77},"Performance for the same model and task can vary greatly across configuration values, so relying on a single or a few configurations can introduce bias and misrepresent the model’s true ability.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]