[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-128524-en":3,"doc-seo-128524-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},128524,687207020761,"Patrick","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Clustering Patients using Longitudinal Data - Master Thesis in ICT for Internet and Multimedia","The management of longitudinal datasets in clinical research, especially under missing data, demands careful statistical decisions. Longitudinal records are crucial for assessing disease development and treatment success across multiple time points. This work addresses missingness in longitudinal height and weight measurements for 3,897 patients aged 0–24 by testing imputation methods, selecting the Mean Expected Growth approach, and deriving age-adjusted BMI trajectories. Clustering is performed with a forgetting-factor-informed Gaussian Mixture Model using gold-standard and imputed-data scenarios.","Master Candidate  \nPargol Golmohammadi  \nStudent ID 2005932  \nAcademic Year 2023/2024  \nMaster Thesis in ICT for Internet and Multimedia  \nClustering Patients using Longitudinal Data  \nSupervisors  \nProf. Saikat Chatterjee KTH University, Sweden  \nProf. Federica Battisti University of Padova, Italy  \nCo-supervisor  \nAshish Kumar Karolinska Institute, Sweden  \nTo my husband and parents  \nAbstract  \nThe management of longitudinal datasets in the context of clinical research, particularly in the presence of missing data, is a complex and diverse task that requires meticulous deliberation. Longitudinal datasets are very important in the context of evaluating disease development and treatment success due to their ability to record information over multiple time points. Nonetheless, the occurrence of missing data might be attributed to a range of factors, including patient dropping out, irregular follow-up, or technical errors. In order to tackle this problem, researchers often use advanced statistical methodologies such as imputation methods, which we have used in this work to handle missing data. In our case, we worked on longitudinal height and weight data of 3897 patients between 0 to 24 years old and the missing data ratio of our dataset was around 35% . As we wanted to get the BMIs of the patients and cluster them, at őrst we replaced these missing data with different imputation approaches, and according to the obtained results, we chose the Mean Expected Growth approach and then calculated the BMIs of the patients. Choosing the best clustering method depends on the nature and distribution of data and the problem deőnition and requirements raised in a project. In this research, the Gaussian Mixture Model (GMM) was selected asthe clustering algorithm duetothe Gaussian distribution of the data. The objective was to comprehend the dynamic changes in patient clusters using a novel forgetting factor approach in the context of longitudinal data to identify age-adjusted BMI growth trajectories. Forgetting factor is an approach used in time-series analysis and forecasting that involves assigning weights to previous data that decrease exponentially with time and analyzes previous observations’ effect on future outcomes. Our dataset had a very high percentage of missing data, therefore we chose to cluster the data in two different ways. In the őrst scenario, we separated the data that did not have missing data, performed clustering on them, and considered it as a gold standard. Then, in the second scenario, we imputed the missing data and performed clustering on the entire dataset. By focusing on early life factors such as gestational smoking, lactation, and preśgestational and gestational BMI control, our őndings contribute additional evidence to the OECD guidance regarding high BMI risksand interventions (World Health Organization, 2016[20]) .  \nSommario  \nContents  \nList of Figures xi  \nList of Tables xiii  \nList of Acronyms xix  \n1 Introduction 1  \n1.1 Challenges ................................ 3  \n1.2 Growth Trajectories of Patients .................... 4  \n1.3 Contributions .............................. 5  \n1.4 Thesis Organization ........................... 5  \n2 Background 7  \n2.1 Chapter overview ............................ 7  \n2.2 Clustering algorithms ......................... 8  \n2.2.1 K-means ............................. 8  \n2.2.2 Gaussian mixture model .................... 9  \n2.2.3 K-medoid ............................ 14  \n2.3 Missing data ............................... 16  \n2.4 Imputation methods .......................... 23  \n2.4.1 Linear interpolation: ...................... 23  \n2.4.2 Forward-Fill(ffill) and Backward-Fill(bőll) .......... 24  \n3 Analysis 27  \n3.1 Dataset description ........................... 27  \n3.2 Mean expected growth ......................... 29  \n3.3 Multivariate imputation by chained equation (MICE) ....... 32  \n3.3.1 The core principles behind multiple imputation (MI) ... 32  \n3.3.2 MI","cbCaiqt143Nk4Vfi","https://ap.wps.com/l/cbCaiqt143Nk4Vfi","pdf",5634681,3,1,99,"English","en",105,"# List of Figures\n# List of Tables\n# List of Acronyms\n# 1 Introduction\n## 1.1 Challenges\n## 1.2 Growth Trajectories of Patients\n## 1.3 Contributions\n## 1.4 Thesis Organization\n# 2 Background\n## 2.1 Chapter overview\n## 2.2 Clustering algorithms\n## 2.3 Missing data\n## 2.4 Imputation methods\n# 3 Analysis\n## 3.1 Dataset description\n## 3.2 Mean expected growth\n## 3.3 Multivariate imputation by chained equation (MICE)\n## 3.4 Cluster selection\n## 3.5 Experimental results\n## 3.6 Results Discussion\n# 4 Conclusions and Future Works","[{\"question\":\"Why is missing data a key challenge in longitudinal clinical datasets?\",\"answer\":\"Missing data can arise from patient dropout, irregular follow-up, or technical issues, and it complicates evaluating disease progression and treatment outcomes across time points.\"},{\"question\":\"What imputation approach was selected for this study and why?\",\"answer\":\"The study compared imputation strategies for longitudinal height and weight data and selected the Mean Expected Growth approach to support subsequent BMI calculation and clustering under high missingness.\"},{\"question\":\"How does the forgetting-factor idea influence patient clustering for BMI trajectories?\",\"answer\":\"Forgetting factor assigns exponentially decreasing weights to older observations, allowing the model to analyze how historical data affects future outcomes while capturing dynamic cluster changes over time.\"}]","Clustering Patients using Longitudinal Data - Master Thesis in ICT for Internet and Multimedia | PDF",1786001547,249,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"clustering-patients-using-longitudinal-data-master-thesis-in-ict-for-internet-and-multimedia","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":20},"https://docshare.wps.com/document/research-report/",{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/clustering-patients-using-longitudinal-data-master-thesis-in-ict-for-internet-and-multimedia/128524/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-23","2026-08-06",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why is missing data a key challenge in longitudinal clinical datasets?","Question",{"text":76,"@type":77},"Missing data can arise from patient dropout, irregular follow-up, or technical issues, and it complicates evaluating disease progression and treatment outcomes across time points.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What imputation approach was selected for this study and why?",{"text":81,"@type":77},"The study compared imputation strategies for longitudinal height and weight data and selected the Mean Expected Growth approach to support subsequent BMI calculation and clustering under high missingness.",{"name":83,"@type":74,"acceptedAnswer":84},"How does the forgetting-factor idea influence patient clustering for BMI trajectories?",{"text":85,"@type":77},"Forgetting factor assigns exponentially decreasing weights to older observations, allowing the model to analyze how historical data affects future outcomes while capturing dynamic cluster changes over time.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]