[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-125537-en":3,"doc-seo-125537-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},125537,1099513958607,"Jiven","https://ap-avatar.wpscdn.com/avatar/100002390cf8733938c?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778829742770036399",8,"Research & Report","Synthetic Dataset Generation for Privacy-Preserving Machine Learning","Machine Learning success depends on large training datasets, yet privacy prevents public release when data contain sensitive information such as medical records. Encryption and obfuscation approaches can protect privacy but often reduce classification accuracy or introduce substantial computational overhead. This paper introduces a method to generate secure synthetic datasets by recording class-wise Batch Normalization statistics from a network pretrained on private data, then optimizing random noise to match layerwise distributions. Experiments on CIFAR10 and ImageNet show comparable training performance from scratch, with high visual dissimilarity and strong privacy under multiple leakage attacks.","Synthetic Dataset Generation for Privacy-Preserving  \nMachine Learning  \nEfstathia Souﬂeri, Gobinda Saha, and Kaushik Roy  \narXiv :2210 .03205v2 [ cs .CR] 10 Oct 2022  \nAbstract—Machine Learning (ML) has achieved enormous success in solving a variety of problems in computer vision, speech recognition, object detection, to name a few. The principal reason for this success is the availability of huge datasets for training deep neural networks (DNNs). However, datasets cannot be publicly released if they contain sensitive information such as medical records, and data privacy becomes a major concern. Encryption methods could be a possible solution, however their deployment on ML applications seriously impacts classiﬁcation accuracy and results in substantial computational overhead. Alternatively, obfuscation techniques could be used, but maintaining a good trade-off between visual privacy and accuracy is challenging. In this paper, we propose a method to generate secure synthetic datasets from the original private datasets. Given a network with Batch Normalization (BN) layers pretrained on the original dataset, we ﬁrst record the class-wise BN layer statistics. Next, we generate the synthetic dataset by optimizing random noise such that the synthetic data match the layerwise statistical distribution of original images. We evaluate our method on image classiﬁcation datasets (CIFAR10, ImageNet) and show that synthetic data can be used in place of the original CIFAR10/ImageNet data for training networks from scratch, producing comparable classiﬁcation performance. Further, to analyze visual privacy provided by our method, we use Image Quality Metrics and show high degree of visual dissimilarity between the original and synthetic images. Moreover, we show that our proposed method preserves data-privacy under various privacy-leakage attacks including Gradient Matching Attack, Model Memorization Attack, and GAN-based Attack.  \nIndex Terms—Synthetic Images, Privacy, Deep Learning, Neural Networks, Privacy-Preserving Machine Learning  \nI. INTRODUCTION  \nMACHINE Learning (ML) has been integrated with  \ngreat success in a wide range of applications such as computer vision, autonomous driving, speech recognition, natural language processing, object detection and so on. The availability of large datasets and advancements in techniques  \nThis work was supported in part by the Center for Brain Inspired Computing (C-BRIC), one of the six centers in Joint University Microelectronics Program (JUMP), in part by the Semiconductor Research Corporation (SRC) Program sponsored by Defense Advanced Research Projects Agency (DARPA), in part by the Semiconductor Research Corporation, in part by the National Science Foundation; in part by the Department of Defense (DoD) Vannevar Bush Fellowship, in part by the U.S. Army Research Laboratory, and in part by IARPA.  \nEfstathia Souﬂeri is with School of Electrical and Computer Engineering, Purdue University, West Lafayette, IN 47907, USA (e-mail: esouﬂ[er@purdue.edu](er@purdue.edu)).  \nGobinda Saha is with School of Electrical and Computer Engineering, Purdue University, West Lafayette, IN 47907, USA (e-mail: [gsaha@purdue.edu](gsaha@purdue.edu)).  \nKaushik Roy is with School of Electrical and Computer Engineering, Purdue University, West Lafayette, IN 47907, USA (e-mail: [kaushik@purdue.edu](kaushik@purdue.edu)).  \n(Under Submission)  \nfor training deep neural network models have played integral roles towards such success. Moreover, the cloud providers offer various Machine Learning as a Service (MLaaS) platforms such as Microsoft Azure ML Studio [1], Google Cloud ML Engine [2], and Amazon Sagemaker [3] etc., where computational resources is provided for running ML workloads. For such cloud-based computing, the ML algorithms are either provided by the users or selected from the standard ML algorithm libraries [4], whereas the datasets are usually shared to cloud by the users to meet application-speciﬁc require","cbCainiRkK3AYdWI","https://ap.wps.com/l/cbCainiRkK3AYdWI","pdf",9010948,1,11,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"Why can’t private datasets always be released for machine learning training?\",\"answer\":\"Because datasets may contain sensitive information such as medical or financial records, which creates serious privacy concerns when shared publicly.\"},{\"question\":\"How does the proposed method generate synthetic datasets?\",\"answer\":\"It records class-wise Batch Normalization layer statistics from a network pretrained on the private dataset, then optimizes random noise so the synthetic data match the original layerwise statistical distributions.\"},{\"question\":\"What evidence is provided that synthetic data maintain both utility and privacy?\",\"answer\":\"The method evaluates on image classification datasets (CIFAR10 and ImageNet), showing comparable performance when training from scratch, and it uses visual quality metrics plus privacy-leakage attacks to demonstrate high dissimilarity and preserved privacy.\"}]","Synthetic Dataset Generation for Privacy-Preserving Machine Learning | PDF",1785899728,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"synthetic-dataset-generation-for-privacy-preserving-machine-learning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/synthetic-dataset-generation-for-privacy-preserving-machine-learning/125537/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why can’t private datasets always be released for machine learning training?","Question",{"text":75,"@type":76},"Because datasets may contain sensitive information such as medical or financial records, which creates serious privacy concerns when shared publicly.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed method generate synthetic datasets?",{"text":80,"@type":76},"It records class-wise Batch Normalization layer statistics from a network pretrained on the private dataset, then optimizes random noise so the synthetic data match the original layerwise statistical distributions.",{"name":82,"@type":73,"acceptedAnswer":83},"What evidence is provided that synthetic data maintain both utility and privacy?",{"text":84,"@type":76},"The method evaluates on image classification datasets (CIFAR10 and ImageNet), showing comparable performance when training from scratch, and it uses visual quality metrics plus privacy-leakage attacks to demonstrate high dissimilarity and preserved privacy.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]