[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120864-en":3,"doc-seo-120864-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120864,1649267921044,"Ava Thompson","https://us-avatar.wpscdn.com/avatar/1800007509477c92dfb?_k=1782875107921204101",8,"Research & Report","Assessment of Differentially Private Synthetic Data for Utility and Fairness in End-to-End Machine Learning Pipelines for Tabular Data - Paper","Differentially private (DP) synthetic data enables sharing while protecting individuals’ privacy. This study examines when synthetic tabular data can substitute for real data in end-to-end machine learning pipelines and identifies effective synthetic data generation methods for model training and evaluation. Utility and fairness are analyzed for downstream classification tasks, considering marginal-based and GAN-based synthetic data algorithms. Results show marginal-based generators outperform GAN-based ones for utility, with MWEM PGM achieving utility and fairness close to models trained on real data.","arXiv :2310 . 19250v1 [ cs .LG] 30 Oct 2023  \nAssessment of Differentially Private Synthetic Data for Utility and Fairness in End-to-End Machine Learning Pipelines for Tabular Data  \nMayana Pereira 1,2,* , Meghana Kshirsagar 1 , Sumit Mukherjee3 , Rahul Dodhia 1 , Juan Lavista Ferres 1 , Rafael de Sousa2  \n1 AI for Good Research Lab, Microsoft, Redmond, Washington, U.S.A.  \n2 Department of Electrical Engineering, University of Brasilia, Brasilia, Brazil  \n3 INSITRO, San Francisco, CA, U.S.A.  \n* [mayana.wanderley@microsoft.com](mayana.wanderley@microsoft.com)  \nAbstract  \nDifferentially private (DP) synthetic data sets are a solution for sharing data while preserving the privacy of individual data providers. Understanding the effects of utilizing DP synthetic data in end-to-end machine learning pipelines impacts areas such as health care and humanitarian action, where data is scarce and regulated by restrictive privacy laws. In this work, we investigate the extent to which synthetic data can replace real, tabular data in machine learning pipelines and identify the most effective synthetic data generation techniques for training and evaluating machine learning models. We systematically investigate the impacts of differentially private synthetic data on downstream classification tasks from the point of view of utility as well as fairness. Our analysis is comprehensive and includes representatives of the two main types of synthetic data generation algorithms: marginal-based and GAN-based.  \nTo the best of our knowledge, our work is the first that: (i) proposes a training and evaluation framework that does not assume that real data is available for testing the utility and fairness of machine learning models trained on synthetic data; (ii) presents the most extensive analysis of synthetic data set generation algorithms in terms of utility and fairness when used for training machine learning models; and (iii)  \nencompasses several different definitions of fairness.  \nOur findings demonstrate that marginal-based synthetic data generators surpass GAN-based ones regarding model training utility for tabular data. Indeed, we show that models trained using data generated by marginal-based algorithms can exhibit similar utility to models trained using real data. Our analysis also reveals that the marginal-based synthetic data generator MWEM PGM can train models that simultaneously achieve utility and fairness characteristics close to those obtained by models trained with real data.  \nIntroduction  \nDifferential privacy (DP) is the standard for privacy-preserving statistical summaries [1] . Companies such as Microsoft [2], Google [3], Apple [4], and government organizations such as the US Census [5], have successfully applied DP in machine learning and data sharing scenarios. The popularity of DP is due to its strong mathematical guarantees. Differential Privacy guarantees privacy by ensuring that the inclusion or exclusion of  \nany particular individual does not significantly change the output distribution of an algorithm.  \nIn areas ranging from health care, humanitarian action, education, and socioeconomic studies, the publication and sharing of data is crucial for informing society and scientific collaboration. However, the disclosure of such data sets can often reveal private, sensitive information. Privacy-preserving data publishing aims at enabling such collaborations while preserving the privacy of individual entries in the data set. Tabular/categorical data about individuals are relevant in many applications, from health care to humanitarian action. Privacy-preserving data publishing for such data can be done in the form of a synthetic data table that has the same schema and similar distributional properties as the real data. The aim here is to release a perturbed version of the original information, so that it can still be used for statistical analysis, but the privacy of individuals in the database is preserved.  \nThe biggest adv","cbCaiuVzU7eYRWXf","https://ap.wps.com/l/cbCaiuVzU7eYRWXf","pdf",1059033,1,21,"English","en",105,"# Abstract\n# Introduction\n## Differential privacy and its guarantees\n## Synthetic data publishing for tabular data\n## Marginal-based vs GAN-based synthetic data generators\n# Key findings and conclusions","[{\"question\":\"What problem does the paper address?\",\"answer\":\"It studies how differentially private synthetic tabular datasets affect the utility and fairness of end-to-end machine learning models, and whether synthetic data can replace real data in such pipelines.\"},{\"question\":\"How does the paper evaluate synthetic data for machine learning?\",\"answer\":\"It investigates downstream classification tasks from the perspective of utility and fairness, using both training and evaluation without assuming access to real data for those metrics.\"},{\"question\":\"Which type of DP synthetic data generator works better in the results?\",\"answer\":\"Marginal-based synthetic data generators outperform GAN-based ones for training utility on tabular data, and MWEM PGM is reported to yield utility and fairness close to real-data-trained models.\"}]","Assessment of Differentially Private Synthetic Data for Utility and Fairness in End-to-End Machine Learning Pipelines for Tabular Data - Paper | PDF",1785732397,53,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"assessment-of-differentially-private-synthetic-data-for-utility-and-fairness-in-end-to-end-machine-learning-pipelines-for-tabular-data-paper","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/assessment-of-differentially-private-synthetic-data-for-utility-and-fairness-in-end-to-end-machine-learning-pipelines-for-tabular-data-paper/120864/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address?","Question",{"text":75,"@type":76},"It studies how differentially private synthetic tabular datasets affect the utility and fairness of end-to-end machine learning models, and whether synthetic data can replace real data in such pipelines.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the paper evaluate synthetic data for machine learning?",{"text":80,"@type":76},"It investigates downstream classification tasks from the perspective of utility and fairness, using both training and evaluation without assuming access to real data for those metrics.",{"name":82,"@type":73,"acceptedAnswer":83},"Which type of DP synthetic data generator works better in the results?",{"text":84,"@type":76},"Marginal-based synthetic data generators outperform GAN-based ones for training utility on tabular data, and MWEM PGM is reported to yield utility and fairness close to real-data-trained models.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]