[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118683-en":3,"doc-seo-118683-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118683,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","Assessing Linkage Risk in Pseudonymized Datasets Under Modern Machine Learning Algorithms - Bachelor’s Thesis","This bachelor’s thesis investigates linkage risk in pseudonymized datasets when modern machine learning algorithms are available and reused. It studies how different pseudonymization techniques, incremental pseudonymization steps, and partial leakage of non-pseudonymized data affect the ability to reidentify individuals across datasets. Experimental results show that even comparatively simple models can achieve strong linkage performance. Stronger cryptographic pseudonymization does not reliably reduce linkage capability, and in some cases pseudonymization may unintentionally facilitate linkage, raising concerns for long-term GDPR-aligned privacy robustness.","Bachelor’s thesis  \nInformation and Communications Technology 2026  \nJuan Sebastian Acosta Der Megerdichian  \nAssessing Linkage Risk in Pseudonymized Datasets Under Modern Machine Learning Algorithms  \nBachelor’s | Abstract  \nTurku University of Applied Sciences Information and Communications Technology 2026 | Total number of pages 38  \nJuan Sebastian Acosta Der Megerdichian  \nAssessing Linkage Risk in Pseudonymized Datasets Under Modern Machine Learning Algorithms  \nAccess to machine learning algorithms has expanded rapidly in recent years, lowering technical barriers and enabling a wider range of actors to perform advanced data analysis. At the same time, large datasets have become increasingly valuable for artificial intelligence training and scientific research, raising significant privacy concerns. In response, the European Union introduced the General Data Protection Regulation (GDPR) in 2016, which promotes pseudonymization as a safeguard to protect individual identities in shared datasets. However, record linkage—the process of identifying the same individuals across multiple datasets—can enable unintended reidentification, particularly when machine learning techniques are applied.  \nThis study adopts an experimental approach to evaluate linkage risk in pseudonymized datasets under multiple conditions. It assesses the performance of several machine learning algorithms across three experimental scenarios: (1) evaluating linkage performance under different pseudonymization techniques,(2) measuring the effects of incremental pseudonymization applied step by step, and (3) testing whether models trained on pseudonymized data can successfully link records when a partially leaked, non-pseudonymizeddataset becomes available.  \nThe results indicate that even relatively simple machine learning models can achieve strong linkage performance across datasets. Increased cryptographic strength in pseudonymization techniques does not consistently correspond to reduced linkage capability, and in some cases pseudonymization appears to  \nsimplify data in ways that facilitate linkage. These results raise concerns about the long-term robustness of current pseudonymization practices in an environment where machine learning tools are widely accessible, datasets are increasingly reused and shared, and data continues to grow in both research and economic value.  \nKeywords: data privacy, record linkage, encryption, pseudonymization, machine learning algorithm  \nContents  \nList of abbreviations 4  \n1 Introduction 5  \n2 Literature Review 8  \n3 Framework 10  \n4 Research methodology 12  \n4.1 Dataset 13  \n4.2Pseudonymization techniques 14  \n4.2.1 Attribute Removal 15  \n4.2.2 Groups of Pseudonymized Attributes 15  \n4.2.3 Evaluated Pseudonymization Methods 16  \n4.2.4 Application Strategy 17  \n4.3Machine Learning Models 17  \n4.3.1 Logistic Regression 17  \n4.3.2 Random Forest Classifier 18  \n4.3.3 Gradient Boosting Classifier 18  \n4.3.4 Deep Neural Network 18  \n4.4Pipelines 19  \n4.4.1 Linear and Neural Model Pipeline 19  \n4.4.2 Tree-Based Model Pipeline 19  \n4.5Experimental Protocol 20  \n5 Results 21  \n5.1 First experiment: Effect of Pseudonymization on GDPR-Protected Columns  \n22  \n5.2Second Experiment: Effect of Incremental Pseudonymization on Linkage Performance 23  \n5.3Third Experiment: Linkage Under Partial Raw Data Leakage 25  \n6 Conclusion and Discussion 26  \n6.1 Interpretation of Experimental Findings 27  \n6.2Implications for Privacy and GDPR Practice 28  \n6.3Limitations and Methodological Considerations 28  \n6.4Future Research Directions 29  \n6.5Overall Conclusion from the Discussion 29  \nReferences 30  \nTables  \nTable 1. Weighted average F1 scores across pseudonymization techniques. 24  \nTable 2. Weighted average F1 scores across pseudonymization steps. 25  \nTable 3. Weighted average F1 scores across increased data linkage. 26  \nList of abbreviations  \nAbbreviation Explanation of abbreviation:  \nAI Artificial Intelligence  \nCSC Finnish IT Cente","cbCaivFYD7yCjUUs","https://ap.wps.com/l/cbCaivFYD7yCjUUs","pdf",469298,1,36,"English","en",105,"# Introduction\n# Literature Review\n# Framework\n# Research methodology\n## Dataset\n## Pseudonymization techniques\n## Machine Learning Models\n## Pipelines\n## Experimental Protocol\n# Results\n## First experiment: Effect of Pseudonymization on GDPR-Protected Columns\n## Second Experiment: Effect of Incremental Pseudonymization on Linkage Performance\n## Third Experiment: Linkage Under Partial Raw Data Leakage\n# Conclusion and Discussion","[{\"question\":\"What privacy risk does the thesis focus on in pseudonymized datasets?\",\"answer\":\"The thesis focuses on record linkage, i.e., the ability to identify the same individuals across multiple datasets, which can lead to unintended reidentification even when pseudonymization is applied.\"},{\"question\":\"Which experimental conditions are used to evaluate linkage risk?\",\"answer\":\"It evaluates linkage performance under different pseudonymization techniques, tests the effects of incremental pseudonymization applied step by step, and checks whether models trained on pseudonymized data can still link records when partially leaked non-pseudonymized data becomes available.\"},{\"question\":\"Do stronger pseudonymization techniques always reduce linkage performance?\",\"answer\":\"No. The results indicate that increased cryptographic strength does not consistently correspond to reduced linkage capability, and pseudonymization may sometimes simplify data in ways that facilitate linkage.\"}]","Assessing Linkage Risk in Pseudonymized Datasets Under Modern Machine Learning Algorithms - Bachelor’s Thesis | PDF",1785684868,91,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"assessing-linkage-risk-in-pseudonymized-datasets-under-modern-machine-learning-algorithms-bachelors-thesis","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/assessing-linkage-risk-in-pseudonymized-datasets-under-modern-machine-learning-algorithms-bachelors-thesis/118683/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04","2026-08-02",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What privacy risk does the thesis focus on in pseudonymized datasets?","Question",{"text":76,"@type":77},"The thesis focuses on record linkage, i.e., the ability to identify the same individuals across multiple datasets, which can lead to unintended reidentification even when pseudonymization is applied.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"Which experimental conditions are used to evaluate linkage risk?",{"text":81,"@type":77},"It evaluates linkage performance under different pseudonymization techniques, tests the effects of incremental pseudonymization applied step by step, and checks whether models trained on pseudonymized data can still link records when partially leaked non-pseudonymized data becomes available.",{"name":83,"@type":74,"acceptedAnswer":84},"Do stronger pseudonymization techniques always reduce linkage performance?",{"text":85,"@type":77},"No. The results indicate that increased cryptographic strength does not consistently correspond to reduced linkage capability, and pseudonymization may sometimes simplify data in ways that facilitate linkage.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]