[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-126713-en":3,"doc-seo-126713-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},126713,962084925782,"Ava Thompson","https://ap-avatar.wpscdn.com/davatar_9964176cb1d06d4a9deccf72a44ae3dc",8,"Research & Report","When Machine Learning Models Leak - An Exploration of Synthetic Training Data","This paper investigates an attack against a machine learning model that predicts whether a person or household will relocate within the next two years using a propensity-to-move classifier. The attacker can query the model for predictions, has public access to marginal distributions of the original training data, and knows non-sensitive attribute values for selected target individuals. The goal is to infer sensitive attributes for those individuals. The study evaluates how training on synthetic data—replacing the original data—affects the attacker’s ability to carry out this sensitive attribute inference.","arXiv :2310 .08775v3 [ cs .LG] 19 May 2024  \nWhen Machine Learning Models Leak: An Exploration of Synthetic Training Data  \nManel Slokom1 ,2 ,3 , Peter-Paul de Wolf ⋆ 2 , and Martha Larson3  \n1 Delft University of Technology, The Netherlands  \n[m.slokom@tudelft.nl](m.slokom@tudelft.nl)  \n2 Statistics Netherlands, The Hague, The Netherlands  \n[pp.dewolf@cbs.nl](pp.dewolf@cbs.nl)  \n3 Radboud University, The Netherlands  \n[m.larson@cs.ru.nl](m.larson@cs.ru.nl)  \nAbstract. We investigate an attack on a machine learning model that predicts whether a person or household will relocate in the next two years, i.e., a propensity-to-move classi􀀌er. The attack assumes that the attacker can query the model to obtain predictions and that the marginal distributions of the data set on which the model was trained are publicly available. The attack also assumes that the attacker has obtained the values of non-sensitive attributes for a certain number of target individuals.  \nThe objective of the attack is to infer the values of sensitive attributes for these target individuals. We explore how replacing the original data with synthetic data when training the model impacts how successfully the attacker can infer sensitive attributes.4  \nKeywords: Synthetic data, model inversion attribute inference attacks, machine learning, propensity to move.  \n1 Introduction  \nGovernmental institutions charged with collecting and disseminating information may use machine learning (ML) models to produce estimates, such as imputing missing values or inferring attributes that cannot be directly observed. When such estimates are published, it is also useful to make the machine learning model itself publicly available, so that researchers using the estimates can evaluate it closely, or even produce their own estimates. Moreover, society also asks for more insight into the models that are used, e.g., to address possible discrimination caused by decisions based on machine learning models.  \n⋆ The views expressed in this paper are those of the authors and do not necessarily re􀀍ect the policy of Statistics Netherlands.  \n4 This paper is a corrected and updated version of the original paper, which was published as: Slokom, M., de Wolf, PP., Larson, M. (2022) . When Machine Learning Models Leak: An Exploration of Synthetic Training Data. In: Domingo-Ferrer, J. , Laurent, M. (eds) Privacy in Statistical Databases. PSD 2022 . Lecture Notes in Computer Science, vol 13463 . Springer, Cham.  \n2 M. Slokom et al.  \nUnfortunately, machine learning models can be attacked in a way that allows an attacker to recover information about the data set that they were trained on [18] . For this reason, making machine learning models available can lead toa risk that information from the training set is leaked. In this paper, we carryout a case study of model inversion attribute inference attacks on a machine learning classi􀀌er to better understand the nature of the risk. Model inversion attribute inference attacks aim to reconstruct the data a model is trained on or expose sensitive information inherent in the data [14,28] . Conventionally, they only seek to infer sensitive attributes of individuals whose data are included in the training set (Inclusive individuals) . Here, we go beyond this conventional perspective to investigate the extent to which the availability of the machine learning model and the marginal distributions of the data it was trained on can support inferring sensitive attributes of individuals who are not in the training set (Exclusive individuals) .  \nThe attack scenario that we study assumes that the classi􀀌er has been made accessible and can be queried with arbitrary input an unlimited number of times, and also that the marginal distributions of the data set the model was trained on have been released. The attacker has a set of non-sensitive attributes of the target individuals including the correct value for the propensity to move attribute for “Inclusive individuals","cbCaicS2eMFNWt9t","https://ap.wps.com/l/cbCaicS2eMFNWt9t","pdf",206948,1,17,"English","en",105,"# 1 Introduction\n# 2 Threat Model","[{\"question\":\"What type of machine learning model is studied in the paper?\",\"answer\":\"The paper studies a propensity-to-move classifier that predicts whether a person or household will relocate within the next two years.\"},{\"question\":\"What assumptions define the attack scenario?\",\"answer\":\"The attacker can query the model unlimited times, uses publicly available marginal distributions from the training dataset, and has values of non-sensitive attributes for target individuals.\"},{\"question\":\"How does training with synthetic data affect the attack risk?\",\"answer\":\"Training on synthetic training data makes the resulting classifier slightly less susceptible, reducing how successfully an attacker can infer sensitive attributes.\"}]","When Machine Learning Models Leak - An Exploration of Synthetic Training Data | PDF",1785934365,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"when-machine-learning-models-leak-an-exploration-of-synthetic-training-data","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/when-machine-learning-models-leak-an-exploration-of-synthetic-training-data/126713/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What type of machine learning model is studied in the paper?","Question",{"text":75,"@type":76},"The paper studies a propensity-to-move classifier that predicts whether a person or household will relocate within the next two years.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What assumptions define the attack scenario?",{"text":80,"@type":76},"The attacker can query the model unlimited times, uses publicly available marginal distributions from the training dataset, and has values of non-sensitive attributes for target individuals.",{"name":82,"@type":73,"acceptedAnswer":83},"How does training with synthetic data affect the attack risk?",{"text":84,"@type":76},"Training on synthetic training data makes the resulting classifier slightly less susceptible, reducing how successfully an attacker can infer sensitive attributes.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]