[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123046-en":3,"doc-seo-123046-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123046,8796095461564,"Liam","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","State-of-the-Art Approaches to Enhancing Privacy Preservation of Machine Learning Datasets - A Survey","Privacy-preserving machine learning (PPML) addresses rising privacy risks created by widespread machine learning deployments across telecommunications, financial technology, and surveillance. The survey studies how adversary capabilities drive distinct threat models, including membership and attribute inference as well as data reconstruction. It reviews defenses that protect confidentiality and integrity of training data in centralized and collaborative learning, covering training-data refinement, privacy-preserving data processing, and risk mitigation. It also analyzes the privacy–utility trade-off using cryptographic methods, Differential Privacy, and Trusted Execution Environments, extending discussion to applied domains.","arXiv :2404 . 16847v1 [ cs .CR] 25 Feb 2024  \nState-of-the-Art Approaches to Enhancing Privacy Preservation of Machine Learning Datasets: A Survey  \nChaoyu Zhang  \nApril 29, 2024  \nAbstract  \nThis paper examines the evolving landscape of machine learning (ML) and its profound impact across various sectors, with a special focus on the emerging field of Privacy-preserving Machine Learning (PPML) . As ML applications become increasingly integral to industries like telecommunications, financial technology, and surveillance, they raise significant privacy concerns, necessitating the development of PPML strategies. The paper highlights the unique challenges in safeguarding privacy within ML frameworks, which stem from the diverse capabilities of potential adversaries, including their ability to infer sensitive information from model outputs or training data. We delve into the spectrum of threat models that characterize adversarial intentions, ranging from membership and attribute inference to data reconstruction. The paper emphasizes the importance of maintaining the confidentiality and integrity of training data, outlining current research efforts that focus on refining training data to minimize privacy-sensitive information and enhancing data processing techniques to uphold privacy. Through a comprehensive analysis of privacy leakage risksand countermeasures in both centralized and collaborative learning settings, this paper aims to provide a thorough understanding of effective strategies for protecting ML training data against privacy intrusions. It explores the balance between data privacy and model utility, shedding light on privacy-preserving techniques that leverage cryptographic methods, Differential Privacy, and Trusted Execution Environments. The discussion extends to the application of these techniques insensitive domains, underscoring the critical role of PPML in ensuring the privacy and security of ML systems.  \n1 Introduction  \nRecent developments in the field of machine learning, especially in the domain of deep learning, have markedly impacted numerous sectors, such as next-g network[69], financial technology, and surveillance systems[13, 12], signaling a transformative phase in these industries. Simultaneously, the surge in machine learning-powered artificial intelligence has brought privacy issues to the forefront. Hence, the notion of Privacy-preserving Machine Learning (PPML) has arisen as a prominent area of interest in both the industrial and academic communities.  \nIt is essential to acknowledge that the task of maintaining privacy in machine learning presents distinct challenges, which vary based on the capabilities of potential adversaries. The growing emphasis on understanding and mitigating information leakage from training data involves considering a spectrum of threat models. These models reflect various levels of adversarial power, including the ability to view model outputs, gain access to model parameters, or use certain iterative optimization techniques. The objectives of these adversaries can vary, such as membership inference, attribute inference, property inference, and data reconstruction [10, 46] . Each of these objectives, as identified in multiple studies, introduces specific difficulties in protecting the integrity and confidentiality of training data within machine learning frameworks.  \nIn the field of PPML, a primary focus is on ensuring that the implemented ML models prevent the escape of confidential information from the training data beyond the trusted boundaries of the data sources. Specifically, during training phase, the main issues of privacy leakage are centered around the handling of data and its computation. Current research addresses these challenges through two key approaches: (i) determining methods to refine or filter the training data with the aim of either minimizing or entirely eradicating any information sensitive to privacy; and (ii) developing techniques for processing ","cbCaipqLT3oL60ad","https://ap.wps.com/l/cbCaipqLT3oL60ad","pdf",477787,1,17,"English","en",105,"# Introduction\n## Privacy leakage and PPML motivation\n## Threat models and adversary goals\n## Core defense directions\n# A Taxonomy of Adversary Goals\n## Membership inference attacks\n## Attribute and property inference\n## Data reconstruction attacks\n# Privacy-Preserving Techniques\n## Cryptography and differential privacy\n## Trusted execution environments (TEE/SGX)","[{\"question\":\"What problems does privacy-preserving machine learning address in ML deployments?\",\"answer\":\"It focuses on preventing sensitive information from training data from escaping trusted boundaries when ML models are trained and used. The main concerns include confidentiality and integrity of training data against privacy leakage attacks.\"},{\"question\":\"Which threat models are discussed for privacy leakage in machine learning?\",\"answer\":\"The survey covers adversary objectives such as membership inference, attribute inference, property inference, and data reconstruction. These models reflect different adversary power levels, including access to outputs, parameters, or optimization processes.\"},{\"question\":\"How do current PPML techniques balance privacy and model utility?\",\"answer\":\"Techniques aim to reduce information leakage while maintaining usable model performance, creating a privacy–utility trade-off. The survey highlights approaches based on cryptographic methods, Differential Privacy, and Trusted Execution Environments, including methods applied to centralized and collaborative learning.\"}]","State-of-the-Art Approaches to Enhancing Privacy Preservation of Machine Learning Datasets - A Survey | PDF",1785814379,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"state-of-the-art-approaches-to-enhancing-privacy-preservation-of-machine-learning-datasets-a-survey","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/state-of-the-art-approaches-to-enhancing-privacy-preservation-of-machine-learning-datasets-a-survey/123046/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problems does privacy-preserving machine learning address in ML deployments?","Question",{"text":75,"@type":76},"It focuses on preventing sensitive information from training data from escaping trusted boundaries when ML models are trained and used. The main concerns include confidentiality and integrity of training data against privacy leakage attacks.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which threat models are discussed for privacy leakage in machine learning?",{"text":80,"@type":76},"The survey covers adversary objectives such as membership inference, attribute inference, property inference, and data reconstruction. These models reflect different adversary power levels, including access to outputs, parameters, or optimization processes.",{"name":82,"@type":73,"acceptedAnswer":83},"How do current PPML techniques balance privacy and model utility?",{"text":84,"@type":76},"Techniques aim to reduce information leakage while maintaining usable model performance, creating a privacy–utility trade-off. The survey highlights approaches based on cryptographic methods, Differential Privacy, and Trusted Execution Environments, including methods applied to centralized and collaborative learning.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]