[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117953-en":3,"doc-seo-117953-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117953,1649267921044,"Ava Thompson","https://us-avatar.wpscdn.com/avatar/1800007509477c92dfb?_k=1782875107921204101",8,"Research & Report","LEAVE-ONE-OUT DISTINGUISHABILITY IN MACHINE LEARNING - Research framework for data memorization and privacy leakage","An analytical framework quantifies how a machine learning algorithm’s output distribution shifts when a small number of data points are included in (or removed from) its training set, formalized as leave-one-out distinguishability (LOOD). LOOD is positioned as a tool for measuring data memorization and information leakage, refining empirical metrics of privacy risk. Gaussian processes model algorithm randomness, and extensive experiments validate LOOD using membership inference attacks, identify high-leakage causes, and design optimal leave-one-out queries to enable training-data reconstruction.","arXiv :2309 . 17310v4 [ cs .LG] 17 Apr 2024  \nLEAVE-ONE-OUT DISTINGUISHABILITY IN MACHINE LEARNING  \nJiayuan Ye†, Anastasia Borovykh‡, Soufiane Hayou†, and Reza Shokri† ∗† National University of Singapore,‡ Imperial College London  \nABSTRACT  \nWe introduce an analytical framework to quantify the changes in a machine learning algorithm’s output distribution following the inclusion of a few data points in its training set, a notion we define as leave-one-out distinguishability (LOOD) .  \nThis is key to measuring data memorization and information leakage as well asthe influence of training data points in machine learning. We illustrate how our method broadens and refines existing empirical measures of memorization and privacy risks associated with training data. We use Gaussian processes to model the randomness of machine learning algorithms, and validate LOOD with extensive empirical analysis of leakage using membership inference attacks. Our analytical framework enables us to investigate the causes of leakage and where the leakage is high. For example, we analyze the influence of activation functions, on data memorization. Additionally, our method allows us to identify queries that disclose the most information about the training data in the leave-one-out setting. We illustrate how optimal queries can be used for accurate reconstruction of training data.1  \n1 INTRODUCTION  \nA key question in interpreting a model involves identifying which members of the training set have influenced the model’s predictions for a particular query (Koh & Liang, 2017; Koh et al., 2019; Pruthi et al., 2020) . Our goal is to understand the impact of the presence of a specific data point in the training set of a model on its predictions. This becomes important, for example, when model’s prediction on a data point is primarily due to its own presence (or that of points similar to it) in the training set-a phenomenon referred to as memorization (Feldman, 2020) and can cause leakage of sensitive information about the training data. This enables an adversary to infer which data points were included in the training set (using membership inference attacks), by observing the model’s predictions (Shokri et al., 2017; Zarifzadeh et al., 2023) . One way to measure such influence, memorization, and information leakage, and thus addressing related questions, is by (re)training models with various combinations of data points. This empirical approach, however, is computationally expensive (Feldman & Zhang, 2020), and does not enable efficient analysis of, for example, what data points are influential, which queries they influence the most, and what properties of model (architectures) can impact their influence. These limitations demand an efficient and constructive modeling approach to influence and leakage analysis. This is the focus of the current paper.  \nWe generalize the above-mentioned closely-related concepts (Section 2), and unify them as the statistical divergence between the output distribution of models, trained using (stochastic) machine learning algorithms, when one or more data points are added to (or removed from) the training set. We refer to this measure as leave-one-out distinguishability (LOOD) . A larger LOOD of a training data S in relation to query data Q suggests a larger potential for information leakage about S, when an adversary observes the prediction of models trained on S and queried at Q. As a special variant of LOOD, if we measure the gap between the mean of prediction distributions over leave-one-out models on a query, the resultant mean distance LOOD recovers existing definitions for influence (function) (Steinhardt et al., 2017; Koh & Liang, 2017; Koh et al., 2019) and memorization (selfinfluence) (Feldman & Zhang, 2020; Feldman, 2020) (i.e., a greater mean distance LOOD suggests a larger influence of training data S on query data Q) .  \n∗Authors AB, SH and RS are ordered alphabetically.  \n1The code is available at [https://github.","cbCaicbExHplBvgL","https://ap.wps.com/l/cbCaicbExHplBvgL","pdf",5977672,1,37,"English","en",105,"# Abstract\n# Introduction\n## Influence, memorization, and leakage\n## Leave-one-out distinguishability (LOOD)\n# Analytical method\n## Gaussian process modeling of output distributions\n# Empirical validation and experiments\n## Membership inference attacks and correlation tests\n## Mean-distance LOOD vs leave-one-out retraining\n# Theoretical explanations and algorithms\n## Most-influenced queries and reconstruction\n## Maximization behavior under NNGP and SGD","[{\"question\":\"What does leave-one-out distinguishability (LOOD) measure in machine learning?\",\"answer\":\"LOOD measures how the output distribution of a model changes when specific data points are added to or removed from the training set in a leave-one-out setting.\"},{\"question\":\"How is LOOD used to assess privacy risks?\",\"answer\":\"A larger LOOD suggests greater potential information leakage when an adversary observes model predictions on query data. The paper validates this link using membership inference attacks in leave-one-out experiments.\"},{\"question\":\"Why are Gaussian processes used in the proposed framework?\",\"answer\":\"Gaussian processes are used to model the randomness of machine learning outputs under querying, enabling accurate and efficient estimation of LOOD without repeated retraining.\"}]","LEAVE-ONE-OUT DISTINGUISHABILITY IN MACHINE LEARNING - Research framework for data memorization and privacy leakage | PDF",1785680514,93,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"leave-one-out-distinguishability-in-machine-learning-research-framework-for-data-memorization-and-privacy-leakage","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/leave-one-out-distinguishability-in-machine-learning-research-framework-for-data-memorization-and-privacy-leakage/117953/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05","2026-08-02",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What does leave-one-out distinguishability (LOOD) measure in machine learning?","Question",{"text":76,"@type":77},"LOOD measures how the output distribution of a model changes when specific data points are added to or removed from the training set in a leave-one-out setting.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How is LOOD used to assess privacy risks?",{"text":81,"@type":77},"A larger LOOD suggests greater potential information leakage when an adversary observes model predictions on query data. The paper validates this link using membership inference attacks in leave-one-out experiments.",{"name":83,"@type":74,"acceptedAnswer":84},"Why are Gaussian processes used in the proposed framework?",{"text":85,"@type":77},"Gaussian processes are used to model the randomness of machine learning outputs under querying, enabling accurate and efficient estimation of LOOD without repeated retraining.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]