[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121956-en":3,"doc-seo-121956-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},121956,137441390410,"Hazel","https://ap-avatar.wpscdn.com/avatar/2000252f4ab5702993?_k=1776741390130283984",8,"Research & Report","Towards Privacy-preserving Machine Learning - Generative Modeling and Discriminative Analysis","The digital era relies on abundant data that fuels machine learning in vision, NLP, speech recognition, and recommendation systems, yet data sharing conflicts with privacy and ethics. Sensitive sources—personal information on mobile devices, confidential medical treatments, and financial records—must satisfy strict legal requirements such as GDPR and HIPAA, even when this slows innovation. Large web-scraped datasets further heighten risks of unintentionally exposing private information and copyrighted material. The dissertation studies three connected directions: privacy-preserving generative modeling for synthetic data with rigorous guarantees, privacy attacks and defenses using membership inference to quantify leakage, and practical applications of DP training for sensitive real-world datasets, especially in medicine.","Towards Privacy-preserving Machine Learning: Generative Modeling and Discriminative Analysis  \nDingfan Chen  \nA dissertation submitted towards the degree Doctor of Engineering Science (Dr.-Ing.)  \nof the Faculty of Mathematics and Computer Science of Saarland University  \nSaarbrücken, 2023.  \nDate of Colloquium: 23.04.2024  \nDean of the Faculty: Prof. Dr. Jürgen Steimle  \nChair of the Committee: Prof. Dr. Isabel Valera  \nReviewers: Prof. Dr. Mario Fritz  \nProf. Dr. Antti Honkela  \nProf. Dr. Catuscia Palamidessi  \nAcademic Assistant: Dr. Raouf Kerkouche  \nDedicated to my family  \nAbstract  \nThe digital era is characterized by the widespread availability of rich data, which has fueled the growth of machine learning applications across diverse fields such as computer vision, natural language processing, speech recognition, and recommendation systems. Nevertheless, data sharing is often at odds with serious privacy and ethical issues. The sensitive nature of much of this data, which includes personal information on mobile devices, confidential medical treatments, and financial records, demands a cautious approach to data sharing. This caution is not just a matter of ethical responsibility but also a legal mandate, with stringent regulations like the General Data Protection Regulation (GDPR) and the Health Insurance Portability and Accountability Act (HIPAA) establishing barriers that, while protective, can also impede the pace of technological progress. Additionally, the growing trend of using largescale, web-scraped datasets to build machine learning models raises serious concerns. This approach, often without proper supervision, can unintentionally include private information and copyrighted content not meant for public use, posing risks of privacy violations and legal complications.  \nThis presents a dilemma: the demand for extensive data to power complex machine learning algorithms conflicts with the need to protect personal privacy and intellectual property rights. Addressing this challenge is critical not only to maintain public trust but also to ensure that the progress in machine learning is sustainable, responsible, and aligned with societal values. To this end, this thesis investigates such privacy risks and seeks out viable solutions that permit data sharing within strict privacy constraints. Specifically, this thesis examines three intertwined perspectives within the realm of data privacy in machine learning: ( 1) privacy-preserving generative modeling, which focuses on generating synthetic data while ensuring rigorous privacy guarantees; (2) privacy attack and defense, dedicated to assessing and understanding the actual privacy risks inherent in machine learning models; as well as (3) applications, which emphasizes the implementation of privacy-preserving training methods on real-world sensitive datasets.  \nFirstly, we explore privacy-preserving generative modeling, with the goal of creating synthetic data that maintains characteristics of the population distribution relevant for particular tasks, while adhering to rigorous privacy guarantee. Such synthetic data can be utilized and analyzed as if it were the real data, thus enabling progress and facilitating reproducible research in sensitive domains. The foundation of our approach is rooted in differentially private (DP) generative modeling. Our advancements involve the development of sanitization protocols dedicated to generative modeling (Chapter 2), the design of a generation framework that reduces the inherent complexity of DP training (Chapter 3), and offering a novel unified perspective that presents a joint design surface for systematic investigation into future advancements in the field (Chapter 4) .  \nSecondly, we delve into privacy attack and defense mechanisms, particularly focusing on real-world simulations of privacy threats. Our work primarily scrutinizes the membership inference attack, which attempts to determine whether a particular data sample was p","cbCaitCVCcYHhxfO","https://ap.wps.com/l/cbCaitCVCcYHhxfO","pdf",20789155,1,226,"English","en",105,"# Abstract\n## Privacy-preserving generative modeling\n## Privacy attack and defense\n## Privacy-centric applications in DP learning","[{\"question\":\"How is differential privacy leveraged for real-world applications?\",\"answer\":\"The thesis adapts analytical and design strategies for DP learning in practical settings, with emphasis on medical data where privacy requirements are especially critical.\"}]","Towards Privacy-preserving Machine Learning - Generative Modeling and Discriminative Analysis | PDF",1785808014,570,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"towards-privacy-preserving-machine-learning-generative-modeling-and-discriminative-analysis","",{"@graph":36,"@context":77},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/towards-privacy-preserving-machine-learning-generative-modeling-and-discriminative-analysis/121956/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"How is differential privacy leveraged for real-world applications?","Question",{"text":75,"@type":76},"The thesis adapts analytical and design strategies for DP learning in practical settings, with emphasis on medical data where privacy requirements are especially critical.","Answer","https://schema.org",{"og:url":52,"og:type":79,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":81,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":98,"slug":130},19,"General","general"]