[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119652-en":3,"doc-seo-119652-105":30,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119652,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","Developing Privacy-Preserving Machine Learning Models for Sensitive Data - Privacy-utility trade-off","This study develops privacy-preserving machine learning models that balance data utility with strong protection for sensitive information, with emphasis on healthcare use cases. Synthetic datasets are generated using multiple privacy techniques, and the impact on data fidelity is measured through statistical modeling and comparative evaluation. Wasserstein-1 distance and related heat maps quantify distributional differences between original and synthetic data. Results indicate that appropriately tuned privacy budgets and methodological adjustments can retain key analytical value while reducing personal information exposure.","Developing Privacy-Preserving Machine Learning Models for Sensitive Data  \nZhe Pang Carlson School of Mangement University of Minnesota  \nMinneapolis, US  \n[pang0121@umn.edu](pang0121@umn.edu)  \nAbstract— This study explores the development of privacyprotecting machine learning models that balance data utility and privacy, especially in sensitive areas such as healthcare. By generating synthetic datasets with various privacy techniques, the study assesses how these methods affect data fidelity. Through statistical modeling and comparative analysis using metrics such as Wasserstein-1 distance and related heat maps, the results show that the optimized synthetic data preserves critical analytical value while protecting personal information. The study concluded that a balance between utility and privacy can be achieved with appropriate privacy budgets and methodological adjustments.  \nKeywords—Privacy-Preserving Machine Learning, Differential Privacy, Synthetic Data Generation  \nIntroduction  \nPrivacy and functionality are two vital but often conflicting priorities in today’s society, especially in sensitive areas such as healthcare and finance. Machine learning helps with prediction and final decision making. However, traditional machine learning models tend to emphasize accuracy at the expense of data confidentiality, raising concerns about potential privacy breaches in these sensitive areas (Xu et al., 2021) . The focus of this study is to address this adverse tradeoff by developing privacy-preserving machine learning models that maintain feature relevance without compromising sensitive information.  \nDifferential privacy has gained public recognition to protect personal data in large data sets. Differential privacy ensures that the removal or addition of a single data point does not significantly affect the overall output of the model, thereby preserving the anonymity of the user (Dwork, n.d.) . However, putting differential privacy into a machine learning framework while maintaining feature correlation remains a challenge tobe solved. Recent advances in generative models, in which denoising deep generative networks, offer promising avenues for synthesizing privacy-preserving datasets while maintaining data utility (Loaiza-Ganem et al., 2022) .  \nIn addition to differential privacy, federated learning has emerged as a viable alternative to centralized data training, enabling collaborative model development without directly sharing data (Kairouz et al., 2019) . This approach is currently being used more in the healthcare sector, where fragmentation of data across agencies often impedes the development of effective predictive models. Moreover, semiparametric methods, known for their flexibility in balancing parametric and nonparametric elements, provide a powerful toolkit for evaluating and enhancing privacy-preserving algorithms (Bi & Shen, 2022) .  \nBuilding on these basic frameworks, this research explores techniques for balancing privacy with data utility and security. By utilizing a generative model approach, this research aims to balance the trade-off between privacy and functionality in machine learning to obtain a security model suitable for sensitive data.  \nI. LITERATURE REVEIW  \nThe dataset used in this study included 349 patient records that contained a combination of categorical and numerical variables. These variables include the type of illness, symptoms such as fever, cough and fatigue, demographic factors such as age and gender, and important health indicators such as blood pressure and cholesterol levels. The main goal is to generate synthetic data that is closely related to the original data set, while incorporating different privacy technologies to ensure the security of personal information.  \nII. MODEL DESCRIPTION  \nSynthetic Data Approaches  \nTo achieve synthetic data generation that protects privacy, I tried several statistical modeling techniques. We evaluated the following ways to ensure privacy complia","cbCairntMGe3zX8a","https://ap.wps.com/l/cbCairntMGe3zX8a","pdf",796803,1,4,"English","en",105,"# Abstract\n# Introduction\n# Literature Review\n# Model Description\n## Synthetic Data Approaches\n## Evaluation of Synthetic Data","[{\"question\":\"What problem does this study address in machine learning for sensitive data?\",\"answer\":\"It tackles the trade-off between model utility and data confidentiality, which can cause privacy risks in sensitive domains like healthcare and finance.\"},{\"question\":\"How does the study generate privacy-preserving synthetic data?\",\"answer\":\"It applies multiple statistical approaches, including multivariate normal modeling, Laplace noise for differential privacy, and classification data perturbation to protect sensitive attributes.\"},{\"question\":\"How is synthetic data quality and privacy impact evaluated?\",\"answer\":\"The study uses Wasserstein-1 distance to compare probability distributions and visual heat maps to assess fidelity, showing that moderate noise can preserve key characteristics while improving privacy.\"}]","Developing Privacy-Preserving Machine Learning Models for Sensitive Data - Privacy-utility trade-off | PDF",1785725475,10,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":28},"developing-privacy-preserving-machine-learning-models-for-sensitive-data-privacy-utility-trade-off","",{"@graph":36,"@context":84},[37,53,67],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":21},"https://docshare.wps.com/document/developing-privacy-preserving-machine-learning-models-for-sensitive-data-privacy-utility-trade-off/119652/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What problem does this study address in machine learning for sensitive data?","Question",{"text":74,"@type":75},"It tackles the trade-off between model utility and data confidentiality, which can cause privacy risks in sensitive domains like healthcare and finance.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does the study generate privacy-preserving synthetic data?",{"text":79,"@type":75},"It applies multiple statistical approaches, including multivariate normal modeling, Laplace noise for differential privacy, and classification data perturbation to protect sensitive attributes.",{"name":81,"@type":72,"acceptedAnswer":82},"How is synthetic data quality and privacy impact evaluated?",{"text":83,"@type":75},"The study uses Wasserstein-1 distance to compare probability distributions and visual heat maps to assess fidelity, showing that moderate noise can preserve key characteristics while improving privacy.","https://schema.org",{"og:url":52,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,127,130,133],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":29,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":29,"slug":132},"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]