[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118565-en":3,"doc-seo-118565-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":20,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118565,962085564381,"Clementine","https://ap-avatar.wpscdn.com/davatar_6f874abed73319feea01a86fa6f0fab8",8,"Research & Report","Data-Driven Delta Machine Learning Models for Improved Extrapolation - Research","The work presents a data-driven method for evaluating a machine learning model’s extrapolation accuracy and for reducing extrapolation error in alloy property prediction. Linear regression is used first to capture general trends, followed by more complex models to learn residual errors (deltas). The approach is tested on aluminum alloys using composition and processing as inputs and yield strength as the target, with careful interpolation versus extrapolation dataset splitting. Results show improved extrapolation over standard models and support pairing with other delta-learning, physics-informed, and active learning strategies to cut iterations.","Data-Driven Delta Machine Learning  \nModels for Improved Extrapolation  \nAdam BIRCHALL a*, Isaac CHANG a , Zidong WANG and Carla BARBATTI, b,c  \na* Brunel University London, BCAST, Uxbridge, United Kingdom b Constellium University Technology Centre, Brunel University London, Uxbridge, United Kingdom  \nc Constellium Technology Center, Parc Economique Centr’alp, Voreppe, France  \nABSTRACT  \nThis work demonstrates a method of testing a machine learning model’s extrapolation accuracy, a capability that is significant to efficiently aid with discovery of improved alloys and presents the application of a pure data-driven method to make steps to reduce this extrapolation error. By using linear models to capture general trends in the data and then the subsequent application of more complex machine learning methods, extrapolation capabilities can be reduced. Being purely data-driven, this type of model can be coupled with other Delta-Machine Learning techniques such as those that utilize physics-domain knowledge, and coupled with active learning methods, with better extrapolation capabilities reducing the number of iterations needed to outperform existing alloys.  \n1. Introduction  \nWhen using machine learning techniques to attempt to improve upon an existing dataset, the extrapolation, and consequentially generalization ability, of a model needs to be a priority for a model to accurately predict beyond what has come before. The prediction of mechanical properties of an alloy given its composition and processing is often a goal of many models, with the aim ofusing such a model to find inputs which result in improved properties.  \nTesting of such models generally demonstrates high accuracies when using standard cross validation techniques, but few fail to demonstrate their poor performance when attempting the aim of accurate predictions of properties beyond the dataset, with models generally being phenomenological, generalizing poorly [1].  \nLinear regression models generalize a great deal and as such often result in poor accuracies, but capture the underlying, linear, relationships in adataset. However, in combination with a higher bias models, linear models can provide the general trend which other models then predict the difference of, resulting in higher accuracies [2].  \nDelta machine learning models are such models where a primary model maps a generalized view of the domain, with a secondary model providing  \npredictions of the errors, or deltas, of the first.  \n2. Experimental procedures  \nUsing data collected from the online database, MatMatch [3], the composition, processing, and yield strength of aluminum alloys was recorded into a database. The composition of the alloys and the temper that each entry received was used to form the input to models with yield strength predictions as the target output.  \nTo test extrapolation capabilities, the top 20% of alloys based on yield strength values were set aside to be used as an extrapolation test set, with the remaining 80% used as an interpolation set. Of this interpolation set, a subsequent split was made of 80% for training data, and 20% for the judgement of interpolation accuracy. Models were then trained upon this interpolation training set and evaluated on both interpolation and extrapolation capabilities using respective datasets.  \nDelta models consisted of an initial, high-generalization model, linear regression, used to make general predictions, the errors of which a subsequent higher bias model would attempt to predict. Through the combination of both general predictions from the first model, and error predictions of the second, a final prediction from the Delta model is made. Linear models, Random Forests, SVRs, Neural Networks, and Delta models were all tested with standard scaling techniques used where appropriate, and all implemented within python utilizing Scikit-learn and modules made to conform to itsAPI [4].  \n3. Results and discussion  \nFigure 1. The prediction of ","cbCair9NCUlPIXk0","https://ap.wps.com/l/cbCair9NCUlPIXk0","pdf",276584,2,1,"English","en",105,"# Introduction\n## Extrapolation challenge and generalization\n## Delta machine learning concept\n# Experimental procedures\n## Data source and target definition\n## Interpolation vs extrapolation split\n## Models and implementation\n# Results and discussion\n## Improved extrapolation performance\n## Extensions and active learning use cases","[{\"question\":\"Why is extrapolation accuracy prioritized when improving machine learning models for alloys?\",\"answer\":\"Extrapolation and generalization determine whether predictions remain accurate beyond the original dataset. The work argues that many models perform well under cross validation but fail when tasked with property prediction outside the data range.\"},{\"question\":\"How do the delta machine learning models reduce extrapolation error?\",\"answer\":\"They combine an initial high-generalization model (linear regression) with a secondary higher-bias model that predicts the residual errors (deltas). The final prediction merges the general trend estimate with the learned error correction.\"},{\"question\":\"How is the dataset split to test interpolation versus extrapolation performance?\",\"answer\":\"The top 20% of alloys by yield strength are reserved for an extrapolation test set, while the remaining 80% form the interpolation set. The interpolation set is further split into 80% training data and 20% for judging interpolation accuracy.\"}]","Data-Driven Delta Machine Learning Models for Improved Extrapolation - Research | PDF",1785684258,5,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"data-driven-delta-machine-learning-models-for-improved-extrapolation-research","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/data-driven-delta-machine-learning-models-for-improved-extrapolation-research/118565/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-09-08","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is extrapolation accuracy prioritized when improving machine learning models for alloys?","Question",{"text":75,"@type":76},"Extrapolation and generalization determine whether predictions remain accurate beyond the original dataset. The work argues that many models perform well under cross validation but fail when tasked with property prediction outside the data range.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do the delta machine learning models reduce extrapolation error?",{"text":80,"@type":76},"They combine an initial high-generalization model (linear regression) with a secondary higher-bias model that predicts the residual errors (deltas). The final prediction merges the general trend estimate with the learned error correction.",{"name":82,"@type":73,"acceptedAnswer":83},"How is the dataset split to test interpolation versus extrapolation performance?",{"text":84,"@type":76},"The top 20% of alloys by yield strength are reserved for an extrapolation test set, while the remaining 80% form the interpolation set. The interpolation set is further split into 80% training data and 20% for judging interpolation accuracy.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":29,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":29,"slug":137},19,"General","general"]