[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120124-en":3,"doc-seo-120124-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120124,8796095461564,"Liam","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Multifidelity linear regression for scientific machine learning from scarce data","Machine learning surrogates for complex engineering systems often suffer when training data are scarce, because generating high-fidelity simulation or experimental samples is expensive and budget-constrained. A new multifidelity linear regression strategy is proposed that leverages available data across fidelities and costs, using an approximate control variate framework to construct multifidelity Monte Carlo estimators. Bias–variance analysis yields guarantees on estimator accuracy and robustness to scarce high-fidelity inputs. Numerical experiments show comparable accuracy to a high-fidelity-only baseline with orders-of-magnitude fewer high-fidelity data requirements.","arXiv :2403 .08627v2 [ stat .ML] 1 Jul 2024  \nMultifidelity linear regression for scientific machine learning from scarce data  \nElizabeth Qian∗†, Dayoung Kang∗, Vignesh Sella‡, Anirban Chaudhuri‡  \nJuly 3, 2024  \nAbstract  \nMachine learning (ML) methods, which fit to data the parameters of a given parameterized model class, have garnered significant interest as potential methods for learning surrogate models for complex engineering systems for which traditional simulation is expensive. However, in many scientific and engineering settings, generating high-fidelity data on which to train ML models is expensive, and the available budget for generating training data is limited, so that high-fidelity training data are scarce. ML models trained on scarce data have high variance, resulting in poor expected generalization performance. We propose a new multifidelity training approach for scientific machine learning via linear regression that exploits the scientific context where data of varying fidelities and costs are available: for example, high-fidelity data maybe generated by an expensive fully resolved physics simulation whereas lower-fidelity data may arise from a cheaper model based on simplifying assumptions. We use the multifidelity data within an approximate control variate framework to define new multifidelity Monte Carlo estimators for linear regression models. We provide bias and variance analysis of our new estimators that guarantee the approach’s accuracy and improved robustness to scarce highfidelity data. Numerical results demonstrate that our multifidelity training approach achieves similar accuracy to the standard high-fidelity only approach with orders-of-magnitude reduced high-fidelity data requirements.  \n1 Introduction  \nScientific and engineering decision-making rely on many-query analyses, in which a predictive model must be evaluated many times at different inputs, initializations, or parameters. Examples include optimization, control, and uncertainty quantification. Traditional high-fidelity models of complex scientific and engineering systems are often prohibitively expensive for this many-query setting, necessitating the development and use of computationally efficient surrogate models. Machine learning (ML) methods, which fit to data the parameters of a given parametrized model class, can learn extremely complex functional relationships from data [35, 44] . A growing body of literature uses such methods to learn surrogate models for engineering systems, for example using linear and kernel regressions [56, 61 , 60 , 50 , 51 , 73 , 74 , 53], as well as nonlinear regressions, for example using  \n∗ School of Aerospace Engineering, Georgia Institute of Technology, Atlanta, Georgia, USA  \n†School of Computational Science and Engineering, Georgia Institute of Technology, Atlanta, Georgia, USA ‡Oden Institute for Computational Engineering and Sciences, University of Texas at Austin, Austin, Texas, USA  \nneural networks [71, 44 , 33 , 3 , 42 , 54 , 46] . These works often take for granted the existence of (or ability to generate) a sufficiently large volume of high-fidelity training data for learning an accurate and robust model. This is a barrier to more widespread adoption of ML surrogate modeling methods in engineering and science because realistic budget constraints in these settings mean that highfidelity data are usually scarce, due to the expense of obtaining data through expensive simulations or experiments. ML surrogate models trained on scarce data are sensitive to quirks of the data set and lead to less accurate and robust predictions [18], limiting trust in the learned models for use in high-consequence engineering and scientific application.  \nOne way to address the data scarcity challenge is through multifidelity scientific machine learning, which seeks to learn models from a combination of scarce high-fidelity data and more abundant lower-cost, lower-fidelity data, which are often available in scien","cbCaimejs4IWIMDR","https://ap.wps.com/l/cbCaimejs4IWIMDR","pdf",1569043,1,28,"English","en",105,"# Introduction\n## Data scarcity in scientific ML surrogates\n## Multifidelity approaches and related work\n## Proposed multifidelity regression framework","[{\"question\":\"Why does scientific machine learning struggle when training data are scarce?\",\"answer\":\"High-fidelity data are costly to generate, so models trained on limited samples exhibit high variance and poor expected generalization, reducing prediction accuracy and robustness.\"},{\"question\":\"What is the key idea of the proposed multifidelity linear regression approach?\",\"answer\":\"It exploits varying-fidelity data by using an approximate control variate framework to build new multifidelity Monte Carlo estimators for linear regression models.\"},{\"question\":\"How do the new estimators perform compared with a high-fidelity-only approach?\",\"answer\":\"The method achieves similar accuracy while requiring orders-of-magnitude fewer high-fidelity training data, improving robustness under scarce high-fidelity availability.\"}]","Multifidelity linear regression for scientific machine learning from scarce data | PDF",1785728324,71,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"multifidelity-linear-regression-for-scientific-machine-learning-from-scarce-data","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/multifidelity-linear-regression-for-scientific-machine-learning-from-scarce-data/120124/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why does scientific machine learning struggle when training data are scarce?","Question",{"text":75,"@type":76},"High-fidelity data are costly to generate, so models trained on limited samples exhibit high variance and poor expected generalization, reducing prediction accuracy and robustness.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the key idea of the proposed multifidelity linear regression approach?",{"text":80,"@type":76},"It exploits varying-fidelity data by using an approximate control variate framework to build new multifidelity Monte Carlo estimators for linear regression models.",{"name":82,"@type":73,"acceptedAnswer":83},"How do the new estimators perform compared with a high-fidelity-only approach?",{"text":84,"@type":76},"The method achieves similar accuracy while requiring orders-of-magnitude fewer high-fidelity training data, improving robustness under scarce high-fidelity availability.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]