[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122853-en":3,"doc-seo-122853-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122853,4398048949847,"Eliana","https://ap-avatar.wpscdn.com/avatar/400002536579ef2da7f?_k=1778318612642679267",8,"Research & Report","Rethinking Privacy in Machine Learning Pipelines - An Information Flow Control Perspective","Modern machine learning relies on models trained over expanding corpora while discarding metadata such as ownership, access control, or licensing during training. Conventional privacy measures like dataset sanitization and differentially private training introduce privacy/utility trade-offs and struggle when sensitive data is shared across participants with fine-grained access needs. This work applies an information flow control viewpoint to incorporate access-control policies, enabling interpretable privacy and confidentiality guarantees and non-interference for user-level settings.","arXiv :2311 . 15792v1 [ cs .LG] 27 Nov 2023  \nRethinking Privacy in Machine Learning Pipelines from an Information Flow Control Perspective  \nLUKAS WUTSCHITZ, M365 Research, UK BORIS KÖPF, Azure Research, UK  \nANDREW PAVERD, Microsoft Security Response Center, UK  \nSARAVAN RAJMOHAN, M365 Research, UK AHMED SALEM, Azure Research, UK SHRUTI TOPLE, Azure Research, UK MENGLIN XIA, M365 Research, UK  \nSANTIAGO ZANELLA-BÉGUELIN, Azure Research, UK VICTOR RÜHLE, M365 Research, UK  \nModern machine learning systems use models trained on ever-growing corpora. Typically, metadata such as ownership, access control, or licensing information is ignored during training. Instead, to mitigate privacy risks, we rely on generic techniques such as dataset sanitization and differentially private model training, with inherent privacy/utility trade-offs that hurt model performance. Moreover, these techniques have limitations in scenarios where sensitive information is shared across multiple participants and fine-grained access control is required. By ignoring metadata, we therefore miss an opportunity to better address security, privacy, and confidentiality challenges.  \nIn this paper, we take an information flow control perspective to describe machine learning systems, which allows us to leverage metadata such as access control policies and define clear-cut privacy and confidentiality guarantees with interpretable information flows. Under this perspective, we contrast two different approaches to achieve user-level non-interference: 1) fine-tuning per-user models, and 2) retrieval augmented models that access user-specific datasets at inference time. We compare these two approaches to a trivially non-interfering zero-shot baseline using a public model and to a baseline that fine-tunes this model on the whole corpus. We evaluate trained models on two datasets of scientific articles and demonstrate that retrieval augmented architectures deliver the best utility, scalability, and flexibility while satisfying strict non-interference guarantees.  \nAdditional Key Words and Phrases: Machine learning, Information flow control, Security, Privacy, Data protection  \n1 INTRODUCTION  \nRecent advances in generative machine learning have been enabled by ever-increasing model sizes and training corpora. These models show impressive performance on many tasks, most prominently language modelling and image generation. This has lead to their wide adoption in real world applications. However, it has also been shown that these models can memorize and subsequently leak information about their training data [Carlini et al. 2023, 2021b] . To remedy this, models are trained on sanitized datasets [Lukas et al. 2023; Zhao et al. 2022] or with anonymization techniques such as differentially private stochastic gradient descent (DP-SGD) [Abadi et al. 2016; Song et al. 2013] . Originally proposed for privacy-preserving statistical databases, the use of differential privacy is not straightforward in other domains such as natural language, as highlighted by Brown et al. [2022] . Moreover, differential privacy is a continuum of guarantees with real-valued parameters 􀁙 and 􀁘 . These parameters are hard to interpret and select, with the debate about what choices are safe for a given application still ongoing.  \n2 L. Wutschitz, B. Köpf, A. Paverd, S. Rajmohan, A. Salem, S. Tople, M. Xia, S. Zanella-Béguelin and V. Rühle  \n(􀀙 public, ⊥)  \n(􀀙, 􀀥dataset)  \n(􀀦, 􀀥required)  \n(􀀤, 􀀥output)  \nFig. 1. Illustration of a machine learning pipeline where inputs and outputs have policies that govern how data can be used. This policy could be a set of users that are allowed to access the data. In this example, the public pre-training dataset 􀀙public has a trivial policy (indicated by ⊥), however, the sensitive training dataset 􀀙 and the query 􀀦 have stricter policies. In the framework we propose, the model produces an output 􀀤 with a policy 􀀥output that is compatible with the policies of all inputs.  ","cbCaieFVkDFucfPh","https://ap.wps.com/l/cbCaieFVkDFucfPh","pdf",673600,1,23,"English","en",105,"# Introduction\n## Motivation\n## Problem\n## Threat Model","[{\"question\":\"Why is privacy at risk in typical machine learning pipelines?\",\"answer\":\"Most systems ignore metadata such as ownership and access control during training, which limits fine-grained protection. Models can also memorize and leak information from training data, making privacy harder to maintain.\"},{\"question\":\"What information flow control perspective contributes to privacy guarantees?\",\"answer\":\"It models machine learning systems using interpretable information flows tied to metadata policies. This enables clear-cut privacy and confidentiality guarantees compatible with the policies of all inputs.\"},{\"question\":\"How do the paper’s user-level non-interference approaches differ?\",\"answer\":\"The paper compares per-user fine-tuning models with retrieval-augmented models that use user-specific datasets at inference time. Experiments show retrieval-augmented architectures provide the best utility, scalability, and flexibility while meeting strict non-interference.\"}]","Rethinking Privacy in Machine Learning Pipelines - An Information Flow Control Perspective | PDF",1785813303,58,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"rethinking-privacy-in-machine-learning-pipelines-an-information-flow-control-perspective","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/rethinking-privacy-in-machine-learning-pipelines-an-information-flow-control-perspective/122853/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is privacy at risk in typical machine learning pipelines?","Question",{"text":75,"@type":76},"Most systems ignore metadata such as ownership and access control during training, which limits fine-grained protection. Models can also memorize and leak information from training data, making privacy harder to maintain.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What information flow control perspective contributes to privacy guarantees?",{"text":80,"@type":76},"It models machine learning systems using interpretable information flows tied to metadata policies. This enables clear-cut privacy and confidentiality guarantees compatible with the policies of all inputs.",{"name":82,"@type":73,"acceptedAnswer":83},"How do the paper’s user-level non-interference approaches differ?",{"text":84,"@type":76},"The paper compares per-user fine-tuning models with retrieval-augmented models that use user-specific datasets at inference time. Experiments show retrieval-augmented architectures provide the best utility, scalability, and flexibility while meeting strict non-interference.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]