[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-126120-en":3,"doc-seo-126120-105":31,"detail-sidebar-cat-0-en-105":97},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},126120,5909887254083,"Miles","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Probing out-of-distribution generalization in machine learning for materials - Research findings summary","Scientific machine learning seeks generalizable models, yet assessment of out-of-distribution (OOD) generalizability often relies on heuristic task definitions. This study systematically evaluates model performance on over 700 OOD tasks in materials datasets, targeting new chemistry or structural symmetry absent from training. Most tasks show strong performance across multiple model classes, including boosted trees. Representation-space analysis links poor performance to test points outside training coverage, where scaling training size or time provides only marginal or adverse gains, indicating many OOD tests probe interpolation rather than true extrapolation.","arXiv :2406 .06489v1 [ cond-mat .mtrl-sci ] 10 Jun 2024  \nProbing out-of-distribution generalization in machine learning for materials  \nKangming Li, 1, ∗ Andre Niyongabo Rubungo,2 Xiangyun Lei,3 Daniel Persaud, 1 Kamal Choudhary,4 Brian DeCost,4 Adji Bousso Dieng,2 and Jason Hattrick-Simpers1, 5, 6, 7,†  \n1 Department of Materials Science and Engineering, University of Toronto, 27 King’s College Cir, Toronto, ON, Canada.  \n2 Vertaix, Department of Computer Science, Princeton University, Princeton, NJ, 08544, USA.  \n3 Toyota Research Institute, 4440 El Camino Real, Los Altos, California 94022, USA.  \n4 Material Measurement Laboratory, National Institute of Standards and Technology, 100 Bureau Dr, Gaithersburg, MD, USA.  \n5 Acceleration Consortium, University of Toronto, 27 King’s College Cir, Toronto, ON, Canada.  \n6 Vector Institute for Artificial Intelligence, 661 University Ave, Toronto, ON, Canada.  \n7 Schwartz Reisman Institute for Technology and Society, 101 College St, Toronto, ON, Canada.  \nScientific machine learning (ML) endeavors to develop generalizable models with broad applicability. However, the assessment of generalizability is often based on heuristics. Here, we demonstrate in the materials science setting that heuristics based evaluations lead to substantially biased conclusions of ML generalizability and benefits of neural scaling. We evaluate generalization performance in over 700 out-of-distribution tasks that features new chemistry or structural symmetry not present in the training data. Surprisingly, good performance is found in most tasks and across various ML models including simple boosted trees. Analysis of the materials representation space reveals that most tasks contain test data that lie in regions well covered by training data, while poorly-performing tasks contain mainly test data outside the training domain. For the latter case, increasing training set size or training time has marginal or even adverse effects on the generalization performance, contrary to what the neural scaling paradigm assumes. Our findings show that most heuristically-defined out-ofdistribution tests are not genuinely difficult and evaluate only the ability to interpolate. Evaluating on such tasks rather than the truly challenging ones can lead to an overestimation of generalizability and benefits of scaling.  \nI. INTRODUCTION  \nMachine learning (ML) has emerged as an important tool in accelerating scientific discovery [1–4] . This transition to datadriven science is epitomized by the development of scientific ML that seeks to build generalizable models capable of broad applicability [5] . In chemical and materials sciences, recent studies have been focused on developing universal or foundational deep learning models, which are suggested to achieve unprecedented levels of out-of-distribution (OOD) generalizations towards unseen materials that are dissimilar to the training data [6–11] .  \nHowever, a critical issue that has been overlooked is the potential biases in selecting OOD tasks to demonstrate generalizability. OOD tasks are often defined based on simple heuristics, which, due to their subjective nature, vary between studies and even lead to contradicting interpretation of generalizability. For instance, the generalization to structures with 5+ elements, despite their omission from training, was used to showcase the emergent capability of deep learning models [6] . However, this perspective has been contested with the argument that such generalizations are anticipated from the physical heuristic that interactions in higher-order systems can be inferred from lower-order ones [12–15] . This discrepancy underscores a broader lack of agreement and discussion on what constitutes a genuinely challenging OOD task. Indeed, if the OOD test set falls within the training domain, conclusions in the superiority of a state-of-the-art ML architecture and the real benefits of model scaling may pertain only to interpolation capabilit","cbCaip2jdkXm2dwv","https://ap.wps.com/l/cbCaip2jdkXm2dwv","pdf",9634994,7,1,31,"English","en",105,"# Introduction\n## Evaluation motivation and bias in OOD task selection\n# Results\n## Evaluation setup","[{\"question\":\"Why are heuristic-defined OOD tasks considered problematic in this study?\",\"answer\":\"OOD tasks are often chosen using simple heuristics that can vary across studies and may place test data inside regions covered by training. This can blur the distinction between interpolation and true extrapolation.\"},{\"question\":\"How many OOD tasks are evaluated, and what are they designed to test?\",\"answer\":\"More than 700 OOD tasks are evaluated, designed to challenge common heuristics by introducing new chemistry elements or structural symmetry groups not present in training data.\"},{\"question\":\"What does representation-space analysis reveal about well-performing vs poorly-performing tasks?\",\"answer\":\"Well-performing tasks’ test data mostly lie within regions well covered by training data, while poorly-performing tasks’ test data fall outside the training domain coverage.\"},{\"question\":\"How does increasing training set size or training time affect generalization on difficult OOD tasks?\",\"answer\":\"For the hardest OOD cases, increasing training size or training time yields marginal improvements or can even degrade generalization performance, contrary to neural scaling expectations.\"}]","Probing out-of-distribution generalization in machine learning for materials - Research findings summary | PDF",1785903269,78,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":92,"head_meta":94,"extra_data":96,"updated_unix":29},"probing-out-of-distribution-generalization-in-machine-learning-for-materials-research-findings-summary","",{"@graph":37,"@context":91},[38,55,70],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":54},"https://docshare.wps.com/document/probing-out-of-distribution-generalization-in-machine-learning-for-materials-research-findings-summary/126120/",4,{"url":53,"name":13,"@type":56,"author":57,"headline":13,"publisher":59,"fileFormat":62,"inLanguage":24,"description":14,"dateModified":63,"datePublished":64,"encodingFormat":62,"isAccessibleForFree":65,"interactionStatistic":66},"DigitalDocument",{"name":9,"@type":58},"Person",{"url":42,"name":60,"@type":61},"DocShare","Organization","application/pdf","2026-08-24","2026-08-05",true,{"@type":67,"interactionType":68,"userInteractionCount":20},"InteractionCounter",{"@type":69},"ViewAction",{"@type":71,"mainEntity":72},"FAQPage",[73,79,83,87],{"name":74,"@type":75,"acceptedAnswer":76},"Why are heuristic-defined OOD tasks considered problematic in this study?","Question",{"text":77,"@type":78},"OOD tasks are often chosen using simple heuristics that can vary across studies and may place test data inside regions covered by training. This can blur the distinction between interpolation and true extrapolation.","Answer",{"name":80,"@type":75,"acceptedAnswer":81},"How many OOD tasks are evaluated, and what are they designed to test?",{"text":82,"@type":78},"More than 700 OOD tasks are evaluated, designed to challenge common heuristics by introducing new chemistry elements or structural symmetry groups not present in training data.",{"name":84,"@type":75,"acceptedAnswer":85},"What does representation-space analysis reveal about well-performing vs poorly-performing tasks?",{"text":86,"@type":78},"Well-performing tasks’ test data mostly lie within regions well covered by training data, while poorly-performing tasks’ test data fall outside the training domain coverage.",{"name":88,"@type":75,"acceptedAnswer":89},"How does increasing training set size or training time affect generalization on difficult OOD tasks?",{"text":90,"@type":78},"For the hardest OOD cases, increasing training size or training time yields marginal improvements or can even degrade generalization performance, contrary to neural scaling expectations.","https://schema.org",{"og:url":53,"og:type":93,"og:title":13,"og:site_name":60,"og:description":14},"article",{"robots":95,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":98},[99,103,107,111,116,121,125,128,133,136,140],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":104,"show_sort_weight":105,"slug":106},"Literature",80,"literature",{"id":54,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},"Exam",70,"exam",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},5,"Comic",60,"comic",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},6,"Technology",50,"technology",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":122,"show_sort_weight":123,"slug":124},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":126,"slug":127},30,"research-report",{"id":129,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":131,"slug":132},9,"Religion & Spirituality",20,"religion-spirituality",{"id":131,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":131,"slug":135},"World Cup","world-cup",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":137,"slug":139},10,"Lifestyle","lifestyle",{"id":141,"doc_module":4,"doc_module_name":47,"category_name":142,"show_sort_weight":112,"slug":143},19,"General","general"]