[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123972-en":3,"doc-seo-123972-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123972,2336464648322,"Aria","https://ap-avatar.wpscdn.com/avatar/2200025388227c56fec?_k=1778556882303663488",8,"Research & Report","Quantifying Distribution Shifts and Uncertainties for Enhanced Model Robustness in Machine Learning Applications","Distribution shifts—mismatches in statistical properties between training and test data—undermine generalization and robustness in real-world machine learning deployments. This study examines model adaptation and generalization using synthetic data to systematically address distributional disparities. Synthetic datasets are generated via the van der Waals equation for gases, and similarity is evaluated with Kullback–Leibler divergence, Jensen–Shannon distance, and Mahalanobis distance. The resulting measures support both accuracy assessment and uncertainty quantification in predictions. Findings indicate Mahalanobis-distance-based identification of low-error interpolation versus high-error extrapolation regimes as a complementary robustness diagnostic. These results strengthen reliability for practical ML deployment under evolving data.","arXiv :2405 .01978v1 [ cs .LG] 3 May 2024  \nQuantifying Distribution Shifts and Uncertainties for Enhanced Model Robustness in Machine Learning Applications  \nVegard Flovik ∗  \nDepartment of Electronic Systems, Norwegian University of Science and Technology,  \nTrondheim, Norway  \nMay 6, 2024  \nAbstract  \nDistribution shifts, where statistical properties differ between training and test datasets, present a significant challenge in real-world machine learning applications where they directly impact model generalization and robustness. In this study, we explore model adaptation and generalization by utilizing synthetic data to systematically address distributional disparities. Our investigation aims to identify the prerequisites for successful model adaptation across diverse data distributions, while quantifying the associated uncertainties. Specifically, we generate synthetic data using the van der Waals equation for gases and employ quantitative measures such as Kullback-Leibler divergence, Jensen-Shannon distance, and Mahalanobis distance to assess data similarity. These metrics enable us to evaluate both model accuracy and quantify the associated uncertainty in predictions arising from data distribution shifts. Our findings suggest that utilizing statistical measures, such as the Mahalanobis distance, to determine whether model predictions fall within the low-error ”interpolation regime” or the high-error ”extrapolation regime” provides a complementary method for assessing distribution shift and model uncertainty. These insights hold significant value for enhancing model robustness and generalization, essential for the successful deployment of machine learning applications in real-world scenarios.  \n1 Introduction  \nIn machine learning, ensuring model accuracy and reliability is paramount for successful deployment in real-world applications. However, challenges arise when the statistical properties of the data differ between training and test datasets, a phenomenon commonly referred to as distribution shift [16] . This presents significant challenges in dynamic systems where data distributions evolve over time, as well as in transfer learning scenarios where models encounter disparate data distributions [15] . Despite efforts to align source and target domains, practical constraints often lead to distributional disparities, impacting model accuracy and reliability in real-world applications [3] .  \nNumerous examples illustrate the implications of distribution shift for machine learning models across various domains. Consider, for example, a scenario where machine learning models monitor the performance of wind turbines in a wind farm. Covariate shift occurs if environmental factors affecting turbine performance, such as wind speed, wind direction, and temperature, differ between historical training data and current operational data due to changes in weather patterns or site conditions. Conversely, target drift arises if the relationship between these environmental factors and turbine performance changes over time due to factors like aging equipment or modifications in operational protocols.  \nSimilarly, in medical image analysis, models trained to detect tumors or diseases from X-rays may fail when deployed in different hospitals with varying equipment, leading to potential misdiagnoses and harmful outcomes for patients [2] . Autonomous vehicles may encounter unforeseen scenarios not adequately represented in training data, posing substantial risks to passenger safety and public trust  \n∗[vflovik@gmail.com](vflovik@gmail.com)  \nin autonomous driving technology [13] . Recommendation systems deployed in online platforms may provide biased or inaccurate suggestions when user preferences evolve over time or when operating in new contexts, leading to sub-optimal user experiences and potential ethical concerns [17] . These examples underscore the critical importance of addressing distribution shift and its impact on model performan","cbCainlIjpVx76hz","https://ap.wps.com/l/cbCainlIjpVx76hz","pdf",1903295,1,15,"English","en",105,"# Abstract\n# 1 Introduction","[{\"question\":\"What problem does the paper address?\",\"answer\":\"The paper addresses distribution shifts, where training and test data have different statistical properties, which can reduce model generalization and robustness.\"},{\"question\":\"How does the study generate synthetic data and compare distributions?\",\"answer\":\"It generates synthetic gas data using the van der Waals equation and compares datasets using Kullback–Leibler divergence, Jensen–Shannon distance, and Mahalanobis distance.\"},{\"question\":\"How do the proposed measures help with uncertainty and robustness?\",\"answer\":\"The metrics are used to evaluate both prediction accuracy and the uncertainty caused by distribution shifts, including distinguishing low-error interpolation from high-error extrapolation regimes using Mahalanobis distance.\"}]","Quantifying Distribution Shifts and Uncertainties for Enhanced Model Robustness in Machine Learning Applications | PDF",1785819501,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"quantifying-distribution-shifts-and-uncertainties-for-enhanced-model-robustness-in-machine-learning-applications","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/quantifying-distribution-shifts-and-uncertainties-for-enhanced-model-robustness-in-machine-learning-applications/123972/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address?","Question",{"text":75,"@type":76},"The paper addresses distribution shifts, where training and test data have different statistical properties, which can reduce model generalization and robustness.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the study generate synthetic data and compare distributions?",{"text":80,"@type":76},"It generates synthetic gas data using the van der Waals equation and compares datasets using Kullback–Leibler divergence, Jensen–Shannon distance, and Mahalanobis distance.",{"name":82,"@type":73,"acceptedAnswer":83},"How do the proposed measures help with uncertainty and robustness?",{"text":84,"@type":76},"The metrics are used to evaluate both prediction accuracy and the uncertainty caused by distribution shifts, including distinguishing low-error interpolation from high-error extrapolation regimes using Mahalanobis distance.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]