[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118838-en":3,"doc-seo-118838-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118838,34359740700684,"Finn","https://ap-avatar.wpscdn.com/avatar/1f400023980c374ae676?_k=1777273430885731487",8,"Research & Report","Privacy Side Channels in Machine Learning Systems - Paper","Most existing privacy defenses for machine learning (ML) treat models as isolated entities, overlooking that real deployments embed multiple system components such as training data filtering, input preprocessing, output monitoring, and query filtering. This work presents privacy side channels: attacks that exploit these components to extract sensitive information at substantially higher rates than standalone-model attacks. Four lifecycle-spanning categories are proposed, including threats that enable membership inference and recovery of users’ private test queries. Results show deduplication can invalidate differential privacy guarantees and that blocked language-model regeneration can still enable exact reconstruction of private keys.","Privacy Side Channels in Machine Learning Systems  \nEdoardo Debenedetti 1 Giorgio Severi2 Nicholas Carlini3 Christopher A. Choquette-Choo3 Matthew Jagielski3 Milad Nasr3 Eric Wallace4 Florian Tramr1  \n1ETH Zurich 2 Northeastern University 3 Google DeepMind 4 UC Berkeley  \narXiv :2309 .05610v1 [ cs .CR] 11 Sep 2023  \nAbstract—Most current approaches for protecting privacy in machine learning (ML) assume that models exist in a vacuum, when in reality, ML models are part of larger systems that include components for training data filtering, output monitoring, and more. In this work, we introduce privacy side channels: attacks that exploit these system-level components to extract private information at far higher rates than is otherwise possible for standalone models. We propose four categories of side channels that span the entire ML lifecycle (training data filtering, input preprocessing, output post-processing, and query filtering) and allow for either enhanced membership inference attacks or even novel threats such as extracting users’ test queries. For example, we show that deduplicating training data before applying differentially-private training creates a side-channel that completely invalidates any provable privacy guarantees. Moreover, we show that systems which block language models from regenerating training data can be exploited to allow exact reconstruction of private keys contained in the training set—even if the model did not memorize these keys. Taken together, our results demonstrate the need for a holistic, end-to-end privacy analysis of machine learning.  \n1. Introduction  \nIn the absence of safeguards, machine learning (ML) models will leak private information about their training data [1, 2] . Numerous methods have been proposed to measure and mitigate these leakages, including formal techniques [3, 4, 5] and heuristics [6, 7] . However, existing methods largely assume that ML models exist in a vacuum, when in reality ML models are part of larger systems that include components for training data filtering, input preprocessing, output monitoring, and more. These systemlevel components are widely incorporated into real-world ML systems to maximize accuracy, security, and robustness.  \nIn this work, we introduce privacy side channels: attacks that exploit system-level components to extract private information at much higher rates than is otherwise possible for isolated ML models. In doing so, we consider adaptive adversaries of varying strengths—ranging from black-box query access to data poisoning capabilities—and show how to conduct privacy attacks that are otherwise impossible without side channels (e.g., revealing test inputs) . Concretely, we propose four categories of attacks that span the entire ML lifecycle (overview in Figure 1):  \n• Training data filtering (Section 3). Most large-scale training sets are filtered to remove duplicates and abnormal examples [6, 7, 8] . We demonstrate that data filters introduce side channels because they create dependencies between different users’ data. In turn, adversaries can amplify privacy attacks by inserting poison examples that maximize these dependencies. Perhaps most surprisingly, we show that data deduplication [6]—a technique designed to improve privacy—can make privacy worse, even causing violations of differential privacy (DP) guarantees (Section 5) . Aside from deduplication, we propose similar attacks for defenses against data poisoning [9, 10, 11] .  \n• Input preprocessing (Section 4). Many models require their inputs to be preprocessed, e.g., language models require text to be tokenized. When these preprocessors are built using training statistics (e.g., tokenizers), we show that it creates side channels that allows adversaries to extract private information such as rare training words [12] .  \n• Model output filtering (Section 4). To improve privacy, many ML systems include filters that prevent models from outputting verbatim training data [13, 14] . We","cbCaik8g9NAq7tTO","https://ap.wps.com/l/cbCaik8g9NAq7tTO","pdf",1858403,1,15,"English","en",105,"# Introduction\n## Training data filtering\n## Input preprocessing\n## Model output filtering\n## Query filtering\n# Conclusion","[{\"question\":\"What are privacy side channels in ML systems?\",\"answer\":\"Privacy side channels are attacks that exploit system-level components around an ML model—such as filtering, preprocessing, output monitoring, and query filtering—to extract private information at much higher rates than attacks on isolated models.\"},{\"question\":\"How do training data filtering and deduplication affect privacy?\",\"answer\":\"Training filters create dependencies across users’ data, letting adversaries amplify privacy attacks by inserting poison examples. The work shows that deduplication, intended to improve privacy, can worsen privacy and even violate differential privacy guarantees.\"},{\"question\":\"What new threats arise from query filtering?\",\"answer\":\"Because query filters aggregate information across users, adversaries can craft targeted inputs to reveal details about other users’ private test queries, which is impossible when analyzing isolated ML models.\"}]","Privacy Side Channels in Machine Learning Systems - Paper | PDF",1785720547,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"privacy-side-channels-in-machine-learning-systems-paper","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/privacy-side-channels-in-machine-learning-systems-paper/118838/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What are privacy side channels in ML systems?","Question",{"text":75,"@type":76},"Privacy side channels are attacks that exploit system-level components around an ML model—such as filtering, preprocessing, output monitoring, and query filtering—to extract private information at much higher rates than attacks on isolated models.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do training data filtering and deduplication affect privacy?",{"text":80,"@type":76},"Training filters create dependencies across users’ data, letting adversaries amplify privacy attacks by inserting poison examples. The work shows that deduplication, intended to improve privacy, can worsen privacy and even violate differential privacy guarantees.",{"name":82,"@type":73,"acceptedAnswer":83},"What new threats arise from query filtering?",{"text":84,"@type":76},"Because query filters aggregate information across users, adversaries can craft targeted inputs to reveal details about other users’ private test queries, which is impossible when analyzing isolated ML models.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]