[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123054-en":3,"doc-seo-123054-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123054,8796095461610,"Oliver","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Improving Realistic Material Property Prediction Using Domain Adaptation Based Machine Learning","Materials property prediction models are commonly evaluated by randomly splitting datasets into training and test sets, which can overestimate performance because material datasets contain substantial redundancy and the split does not match real scientific usage. In practice, researchers predict properties for a known subset of related out-of-distribution (OOD) materials with a different distribution from training data, and they often already know the target material’s formulas/structures. This work proposes domain adaptation to incorporate target material information into learning, improving OOD prediction across five realistic application scenarios. Benchmarks identify DA variants that improve OOD performance, while standard ML and most alternative DA methods do not.","arXiv :2308 .02937v3 [ cond-mat .mtrl-sci ] 27 May 2024  \nIMPROVING REALISTIC MATERIAL PROPERTY PREDICTION USING DOMAIN ADAPTATION BASED MACHINE LEARNING ∗  \nJeffrey Hu  \nDutch Fork High School, Irmo, SC, 39063 Department of Computer Science and Engineering University of South Carolina Columbia, SC 29201  \nDavid Liu  \nDepartment of Electrical Engineering and Computer Science University of Michigan  \nAnn Arbor, MI 48103  \nNihang Fu  \nDepartment of Computer Science and Engineering University of South Carolina Columbia, SC 29201  \nRongzhi Dong  \nDepartment of Computer Science and Engineering University of South Carolina Columbia, SC 29201  \nABSTRACT  \nMaterials property prediction models are usually evaluated using random splitting of datasets into training and test datasets, which not only leads to over-estimated performance due to inherent redundancy, typically existent in material datasets, but also deviate away from the common practice of materials scientists: they are usually interested in predicting properties for a known subset of related out-of-distribution (OOD) materials rather than a universally distributed samples. Feeding such target material formulas/structures to the machine learning models should improve the prediction performance while most current machine learning (ML) models neglect this information. Here we propose to use domain adaptation (DA) to enhance current ML models for property prediction and evaluate their performance improvements in a set of five realistic application scenarios. Our systematic benchmark studies show that there exist DA models that can significantly improve the OOD test set prediction performance while standard ML models and most of the other DAs cannot improve or even deteriorate the performance. Our benchmark datasets and DA code can be freely accessed at [https://github.com/Little-Cheryl/MatDA](https://github.com/Little-Cheryl/MatDA).  \nKeywords material property prediction · out of distribution · domain adaptation · domain shift · machine learning  \n1 Introduction  \nNowadays, machine learning (ML) models are being widely used in materials property prediction for discovering novel materials such as super-hard materials [1, 2], wide band gap materials [3], and energy materials [4] . A large number of innovations have been proposed to improve the ML performance for materials property prediction, including more expressive descriptors [5], better deep learning models (IRNET)[6], graph neural networks that better capture interatomic interactions [7, 8, 9, 10], data augmentation [11], multi-fidelity datasets that combine computational and experimental data [12], active learning[13, 3], and transfer learning[14] . These models and algorithms have significantly improved prediction performance over the past few years. However, it has been found that existing ML algorithms have low generalization performance for test samples with different data distributions, and their prediction performance is often over-estimated due to the high dataset redundancy [15] as many materials are accumulated as a result of a tinkering material discovery process over history. Previously the ML-based material property prediction performances are all evaluated by randomly splitting the whole dataset into training and testing sets. The resulting test set does not share a high degree of homogeneity in terms of composition, structure, or properties, but is randomly distributed in the whole dataset space. This practice does not reflect the realistic application scenario for these ML models when they  \n∗  Citation: Jeffrey Hu et al.. MaterialDA. 15 Pages.... DOI:000000/11111.  \nDomain adaptation for material property prediction  \nare more likely to be applied to predict the properties for a set of similar materials that have a different distribution from the training set, and are located in the sparse chemical space with few known materials, or tend to have extreme property values. Moreover, current ML and deep learning","cbCaijUOdJZrGQa0","https://ap.wps.com/l/cbCaijUOdJZrGQa0","pdf",1858982,1,15,"English","en",105,"# Introduction\n## Motivation: realistic OOD evaluation in materials science\n## Related work on OOD generalization and distribution shift\n## Domain adaptation approaches for OOD prediction","[{\"question\":\"Why do random train-test splits overestimate performance in materials property prediction?\",\"answer\":\"Because materials datasets are highly redundant and the test set is randomly distributed, the evaluation does not reflect realistic application scenarios for distribution-shifted materials.\"},{\"question\":\"How does the proposed approach improve predictions for out-of-distribution materials?\",\"answer\":\"It uses domain adaptation so the model can incorporate target material formulas/structures, aiming to better handle domain shift between training and OOD target materials.\"},{\"question\":\"What do the benchmark studies conclude about domain adaptation versus standard machine learning?\",\"answer\":\"Some domain adaptation models significantly improve OOD test set prediction performance, while standard ML models and most other DA methods cannot improve or may even deteriorate performance.\"}]","Improving Realistic Material Property Prediction Using Domain Adaptation Based Machine Learning | PDF",1785814420,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"improving-realistic-material-property-prediction-using-domain-adaptation-based-machine-learning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/improving-realistic-material-property-prediction-using-domain-adaptation-based-machine-learning/123054/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do random train-test splits overestimate performance in materials property prediction?","Question",{"text":75,"@type":76},"Because materials datasets are highly redundant and the test set is randomly distributed, the evaluation does not reflect realistic application scenarios for distribution-shifted materials.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed approach improve predictions for out-of-distribution materials?",{"text":80,"@type":76},"It uses domain adaptation so the model can incorporate target material formulas/structures, aiming to better handle domain shift between training and OOD target materials.",{"name":82,"@type":73,"acceptedAnswer":83},"What do the benchmark studies conclude about domain adaptation versus standard machine learning?",{"text":84,"@type":76},"Some domain adaptation models significantly improve OOD test set prediction performance, while standard ML models and most other DA methods cannot improve or may even deteriorate performance.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]