[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-128566-en":3,"doc-seo-128566-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},128566,549768064622,"Anda","https://ap-avatar.wpscdn.com/davatar_6f874abed73319feea01a86fa6f0fab8",8,"Research & Report","Exposome Data Drift: Implications for Machine Learning Based Diabetes Prediction","Data drift is a machine learning failure mode where input predictor characteristics change over time, degrading model performance. Yet the impact of data drift on models trained with human exposome data remains insufficiently characterized. This thesis investigates drift in machine learning models for diabetes risk using UK Biobank participants across two time periods, showing significant degradation on the follow-up cohort and identifying covariate, label, and concept drift signatures through drift detection tests.","UNIVERSITAT DE BARCELONA  \nFUNDAMENTAL PRINCIPLES OF DATA SCIENCE MASTER ’S  \nTHESIS  \nExposome Data Drift: Implications for Machine Learning Based Diabetes  \nPrediction  \nAuthor:  \nPeter Hannagan BROSTEN  \nSupervisor: Dr. Karim LEKADIR Marina CAMACHO  \nA thesis submitted in partial fulfillment of the requirements for the degree of MSc in Fundamental Principles of Data Science  \nin the  \nFacultat de Matemàtiques i Informàtica  \nJune 30, 2023  \niii  \nUNIVERSITAT DE BARCELONA  \nAbstract  \nFacultat de Matemàtiques i Informàtica  \nMSc  \nExposome Data Drift: Implications for Machine Learning Based Diabetes  \nPrediction  \nby Peter Hannagan BROSTEN  \nData drift is a problem in machine learning (ML) where characteristics of the input predictors changes over time, leading to model degradation. However, the effects of data drift on ML models built from human exposome data have not been well described yet. This study aimed to investigate data drifts for exposome data in ML models of diabetes risk. 7,521 participants with a diagnosis of diabetes from the UK Biobank, along with a proportional control group from 2006 to 2010 were used to train several baseline ML models for diabetes prediction. A second cohort of 4,007 participants attending the follow-up assessment period from 2012 to 2013 was used to assess potential data drifts over time. When evaluated on the second cohort, significant performance degradation was found in all baseline models (i.e. average precision dropped by 15%, f1-score by 12%, recall by 15%, and precision by 10%) . A suite of drift detection tests were run on the best performing baseline models to identify possible signatures of three distinct kinds of data drift: covariate drift, label drift, and concept drift. Utilizing both multivariate and univariate datadistribution based detection methods, covariate drift was identified in features such as Birth Year, BMI, Frequency of Tiredness, and Lack of Education. A comparison of prevalence rates for time-ordered batches of the population found no severe label drift. Nonetheless, gradual label drift could not be ruled out. A model-aware concept drift detection method was employed, monitoring temporal changes in normalized Shapley contributions for the model’s input features. This test found drift in abnormal changes in feature contribution when predicting on the second cohort for the Birth Year feature and near alerts in multiple others. This study shows the potential for data drift acting as a driver of model degradation in exposome-based ML models and highlights the need for further research into the traceability of clinical AI/ML solutions.  \nv  \nAcknowledgements  \nThis thesis represents the culmination of a year long effort. This was a journey which could not have happened without the help and support of many along the way.  \nI would like to thank the European Union’s Horizon 2020 research and innovation program which funded this research under Grant Agreement Nº 848158 (earlycause.europescience.eu) and the Grant Agreement Nº 874739 ([www.longitools.org](www.longitools.org)); the good people at the BCN-AIM research group, for welcoming and helping me over this last year; and my supervisors Karim Lekadir and Marina Camacho, your wisdom and counsel has been invaluable during this project.  \nTo my parents and sister, thank you for always finding the time to call despite the inconvenience of living on opposite sides of the world. To my friends, Joe and Jessamyn, thank you for listening to me ramble and finding new ways to make me laugh.  \nLastly, to my life partner Elena. Your love is what brought me here and is what helped me preserver. You supported me when I was low and cheered me on when I felt on top of the world. I can never put into words what you mean to me. Thankyou for everything you are.  \nvii  \nContents  \nAbstract iii  \nAcknowledgements v  \n1 Introduction 1  \n2 Preliminaries 3  \n2.1 Probabilistic Underpinnings ......................... 3  \n2.1.1 Random Variabl","cbCaiqHRPCBzOyj8","https://ap.wps.com/l/cbCaiqHRPCBzOyj8","pdf",3037125,2,1,59,"English","en",105,"# Introduction\n# Preliminaries\n## Probabilistic Underpinnings\n## Comparing Distributions\n## Data Drift\n## Drift Detection\n# The Data\n## UKBB Exposome Features\n## Cohort Selection\n## Data Cleaning\n# Model Development\n## Task Definition\n## Model Architectures\n## The Learning Scheme\n# Performance Results","[{\"question\":\"What is data drift, and why does it matter for diabetes prediction models?\",\"answer\":\"Data drift occurs when the characteristics of input predictors change over time, causing model performance to degrade. In exposome-based diabetes prediction, drift can reduce the reliability of models when applied to later cohorts.\"},{\"question\":\"Which datasets and time periods were used to train and evaluate the models?\",\"answer\":\"Models were trained using 7,521 UK Biobank participants with diabetes diagnosis and a proportional control group from 2006 to 2010. Potential drift was then assessed on a second cohort of 4,007 participants from follow-up assessments in 2012 to 2013.\"},{\"question\":\"What types of drift were investigated, and what did the tests find?\",\"answer\":\"The study evaluated covariate drift, label drift, and concept drift. Covariate drift was detected in features such as Birth Year, BMI, Frequency of Tiredness, and Lack of Education, while severe label drift was not observed; gradual label drift could not be ruled out.\"}]","Exposome Data Drift: Implications for Machine Learning Based Diabetes Prediction | PDF",1786001778,149,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"exposome-data-drift-implications-for-machine-learning-based-diabetes-prediction","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,48,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":20},"https://docshare.wps.com/document/","Document",{"item":49,"name":12,"@type":44,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/exposome-data-drift-implications-for-machine-learning-based-diabetes-prediction/128566/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-23","2026-08-06",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What is data drift, and why does it matter for diabetes prediction models?","Question",{"text":76,"@type":77},"Data drift occurs when the characteristics of input predictors change over time, causing model performance to degrade. In exposome-based diabetes prediction, drift can reduce the reliability of models when applied to later cohorts.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"Which datasets and time periods were used to train and evaluate the models?",{"text":81,"@type":77},"Models were trained using 7,521 UK Biobank participants with diabetes diagnosis and a proportional control group from 2006 to 2010. Potential drift was then assessed on a second cohort of 4,007 participants from follow-up assessments in 2012 to 2013.",{"name":83,"@type":74,"acceptedAnswer":84},"What types of drift were investigated, and what did the tests find?",{"text":85,"@type":77},"The study evaluated covariate drift, label drift, and concept drift. Covariate drift was detected in features such as Birth Year, BMI, Frequency of Tiredness, and Lack of Education, while severe label drift was not observed; gradual label drift could not be ruled out.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]