[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119555-en":3,"doc-seo-119555-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119555,1374391974585,"Genevieve","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Differentially Private Machine Learning - Implementation and Analysis of Gradient and Dataset Perturbation Techniques","The increasing use of machine learning poses significant privacy risks, especially when sensitive data is used, and conventional anonymization methods have proven insufficient. Differential privacy is a rigorous framework for data privacy providing strong mathematical guarantees. Applying this framework to machine learning addresses the privacy problem by integrating differential privacy into training pipelines through dataset and gradient perturbation. The study empirically investigates implementations, evaluates privacy-utility trade-offs, and demonstrates feasibility on real-world medical data.","Treball final de grau  \nDOBLE GRAU EN MATEMÀTIQUES IENGINYERIA INFORMÀTICA  \nFacultat de Matemàtiques i Informàtica Universitat de Barcelona  \nDifferentially Private Machine Learning: Implementation and Analysis of Gradient and Dataset Perturbation Techniques  \nAutor: Juan Pablo Mantilla Carreño  \nDirector: Dr. Nahuel Statuto  \nRealitzat a: Departament de Matemàtiques i Informàtica  \nBarcelona, June 10, 2025  \nContents  \nIntroduction iv  \n1 Differential Privacy 1  \n1.1 Background ...................................... 1  \n1.2 Pure Differential Privacy ............................... 2  \n1.3 What is Privacy? ................................... 3  \n1.4 Laplace Mechanism .................................. 4  \n1.5 Properties of Pure Differential Privacy ....................... 6  \n1.6 Approximate Differential Privacy .......................... 7  \n1.7 Gaussian Mechanism ................................. 8  \n1.8 Properties of Approximate Differential Privacy .................. 11  \n2 Machine Learning 17  \n2.1 Core Concepts and Empirical Risk Minimization (ERM) ............ 17  \n2.2 Gradient-Based Optimization for ERM ...................... 18  \n2.3 Model Evaluation and Generalization ....................... 20  \n2.4 Privacy Risks in Standard Machine Learning ................... 21  \n3 Machine Learning with Differential Privacy 23  \n3.1 Gradient Perturbation ................................ 23  \n3.2 Dataset Perturbation ................................. 25  \n4 Empirical Evaluation of Differentially Private Machine Learning Techniques 27  \n4.1 Dataset Description and Origin of Data ...................... 27  \n4.2 Data Preparation and Preprocessing ........................ 28  \n4.3 Implementation of Machine Learning Models and Differential Privacy Mechanism .......................................... 29  \n4.3.1 Baseline Model Implementation ...................... 29  \n4.3.2 Dataset Perturbation ............................. 30  \n4.3.3 Gradient Perturbation: Differentially Private Stochastic Gradient Descent (DP-SGD) ................................ 30  \n4.4 Experimental Results and Analysis ......................... 31  \n4.4.1 Evaluation Metrics and Experimental Setup ............... 31  \n4.4.2 Analysis of Dataset Perturbation Models ................. 32  \n4.4.3 Analysis of Gradient Perturbation Models ................ 34  \n4.4.4 Comparative Analysis and Privacy-Utility Trade-off .......... 36  \nConclusions 37  \nBibliography 39  \nAbstract  \nThe increasing use of machine learning poses significant privacy risks, especially when sensitive data is used, and conventional anonymization methods have proven insufficient. Differential privacy is a rigorous framework for data privacy providing strong mathematical guarantees. The possibility of applying this framework to machine learning solves the privacy problem. We will present the fundamental basis of these concepts to empirically investigate, implement, and analyse two techniques for integrating differential privacy into machine learning pipelines. The first technique, dataset perturbation, involves adding calibrated Gaussian noise directly to the training data and then using any standard machine learning pipeline. The second, gradient perturbation, centers on differentially private stochastic gradient descent, is an approach that injects noise into the gradients during the training phase. For the comparative study, we developed a multi-class classification architecture using a real-world, sensitive medical dataset derived from the MIMIC-IV database. Model performance was evaluated against a non-private baseline, using the appropriate metrics considering our class imbalance, such as Macro F1-score and Macro OVO AUC. The results confirm the trade-off between privacy and utility in the models developed, where higher privacy guarantees consistently result in reduced model utility. For the specific context of this study, gradient perturbation provided a slightly more advantageous model in overa","cbCairBRyuhIFL2t","https://ap.wps.com/l/cbCairBRyuhIFL2t","pdf",1404826,1,49,"English","en",105,"# Introduction\n# Differential Privacy\n## Background\n## Pure Differential Privacy\n## What is Privacy?\n## Laplace Mechanism\n## Properties of Pure Differential Privacy\n## Approximate Differential Privacy\n## Gaussian Mechanism\n## Properties of Approximate Differential Privacy\n# Machine Learning\n## Core Concepts and Empirical Risk Minimization (ERM)\n## Gradient-Based Optimization for ERM\n## Model Evaluation and Generalization\n## Privacy Risks in Standard Machine Learning\n# Machine Learning with Differential Privacy\n## Gradient Perturbation\n## Dataset Perturbation\n# Empirical Evaluation of Differentially Private Machine Learning Techniques\n## Dataset Description and Origin of Data\n## Data Preparation and Preprocessing\n## Implementation of Machine Learning Models and Differential Privacy Mechanism\n## Experimental Results and Analysis\n# Conclusions","[{\"question\":\"What problem does differential privacy address in machine learning?\",\"answer\":\"It provides rigorous mathematical guarantees that limit how much sensitive information can be inferred from training data, countering weaknesses of conventional anonymization methods.\"},{\"question\":\"How do dataset perturbation and gradient perturbation differ?\",\"answer\":\"Dataset perturbation adds calibrated Gaussian noise to the training data before running a standard machine learning pipeline. Gradient perturbation injects noise into gradients during training via differentially private stochastic gradient descent (DP-SGD).\"},{\"question\":\"What do the experimental results show about privacy and utility?\",\"answer\":\"The results confirm a privacy-utility trade-off: higher privacy guarantees consistently reduce model utility. In this study, gradient perturbation achieved a slightly more favorable balance overall.\"}]","Differentially Private Machine Learning - Implementation and Analysis of Gradient and Dataset Perturbation Techniques | PDF",1785724953,123,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"differentially-private-machine-learning-implementation-and-analysis-of-gradient-and-dataset-perturbation-techniques","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/differentially-private-machine-learning-implementation-and-analysis-of-gradient-and-dataset-perturbation-techniques/119555/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does differential privacy address in machine learning?","Question",{"text":75,"@type":76},"It provides rigorous mathematical guarantees that limit how much sensitive information can be inferred from training data, countering weaknesses of conventional anonymization methods.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do dataset perturbation and gradient perturbation differ?",{"text":80,"@type":76},"Dataset perturbation adds calibrated Gaussian noise to the training data before running a standard machine learning pipeline. Gradient perturbation injects noise into gradients during training via differentially private stochastic gradient descent (DP-SGD).",{"name":82,"@type":73,"acceptedAnswer":83},"What do the experimental results show about privacy and utility?",{"text":84,"@type":76},"The results confirm a privacy-utility trade-off: higher privacy guarantees consistently reduce model utility. In this study, gradient perturbation achieved a slightly more favorable balance overall.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]