[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118617-en":3,"doc-seo-118617-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118617,687197207919,"Theodora","https://ap-avatar.wpscdn.com/avatar/a000253d6f5f7c60be?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779446848396160552",8,"Research & Report","Information Flow Control in Machine Learning through Modular Model Architecture","Machine learning models can unintentionally leak sensitive information present in training data via their inference-time outputs, creating a privacy and security risk when users have access to different subsets of data. The work proposes information flow control for machine learning and introduces a Transformer extension that enforces a formal IFC definition. By constraining each security domain’s training influence to a single expert module and enabling only allowed experts at inference under a runtime access policy, the architecture prevents leakage from inaccessible domains. Evaluations on large text and code datasets show low overhead and substantial accuracy gains.","Information Flow Control in Machine Learning through Modular Model Architecture  \narXiv :2306 .03235v2 [ cs .LG] 2 Jul 2024  \nTrishita Tiwari Cornell University  \nSuchin Gururangan University of Washington  \nChuan Guo FAIR at Meta  \nSanjay Kariyappa Georgia Institute of Technology  \nUdit Gupta Cornell University  \nWeizhe Hua ∗  \nGoogle DeepMind Wenjie Xiong Virginia Tech  \nKiwan Maeng Pennsylvania State University  \nHsien-Hsin S. Lee † G. Edward Suh †  \nIntel NVIDIA / Cornell University  \nAbstract  \nIn today’s machine learning (ML) models, any part of the training data can affect the model output. This lack of control for information flow from training data to model output is a major obstacle in training models on sensitive data when access control only allows individual users to access a subset of data. To enable secure machine learning for accesscontrolled data, we propose the notion of information flow control for machine learning, and develop an extension to the Transformer language model architecture that strictly adheres to the IFC definition we propose. Our architecture controls information flow by limiting the influence of training data from each security domain to a single expert module, and only enables a subset of experts at inference time based on the access control policy. The evaluation using large text and code datasets show that our proposed parametric IFC architecture has minimal (1.9%) performance overhead and can significantly improve model accuracy (by 38% for the text dataset, and between 44%–62% for the code datasets) by enabling training on access-controlled data.  \n1 Introduction  \nRecent studies [9, 10, 48] showed that large machine learning (ML) models can leak sensitive information that was in their training data through their inference-time output. This inference-time leakage introduces a new security/privacy concern when different users have access to different (training) data. For example, consider a scenario where a company trains an in-house code-completion model [21] with its proprietary code repositories, when individual employees have access to different subsets of the repositories based on their team or projects. The company faces a dilemma—if the model is trained with all the repositories, employees may be able to learn about the code that they do not have permission to access; if only trained with repositories that everybody has access to, the model quality will degrade. It would be ideal  \n∗ Work done while at Cornell University.†Work done while at Meta.  \nif the model can generate different outputs to different employees, by selectively leveraging training data that they have access to. Unfortunately, today’s ML models mix information from all training data during the training process, and cannot selectively utilize a subset of training data.  \nTo address this problem, we introduce the notion of information flow control (IFC) to machine learning. Instead of treating all training data uniformly, we partition the training dataset into multiple security domains. During inference, a trained ML model takes an access policy that represents which security domains are accessible by the current user and ensures its output does not leak any training data that the user cannot access. The goal of IFC in this ML setting is to develop a training process, an inference process, and a model architecture that can honor the access policy at inference time by guaranteeing no leakage of information from inaccessible training data to the model output. We define the problem through the lens of non-interference (NI), and present IFC asa new challenge that ML systems should address.  \nDeep learning models are notoriously difficult to understand or control even though they perform well in practice. In that sense, IFC in parametric deep neural networks may seem infeasible at a glance. We show that information flow can be controlled at the model architecture level. Our approach trains a separate, small sub-module, ","cbCairrfrjxO3WQi","https://ap.wps.com/l/cbCairrfrjxO3WQi","pdf",1683873,1,18,"English","en",105,"# Abstract\n# Introduction\n## Problem: inference-time leakage under access control\n## Proposed approach: IFC with security domains and access policies\n## Modular architecture with secure gating and expert aggregation\n## Implementation: secure Transformer extensions","[{\"question\":\"What security problem does the paper address in machine learning?\",\"answer\":\"It addresses how ML models can leak sensitive information from training data through inference-time outputs, especially when different users can access different subsets of training data.\"},{\"question\":\"How does the proposed IFC architecture limit information flow?\",\"answer\":\"Training data from each security domain is restricted to influence only a single expert module, and at inference time only the experts permitted by the user’s access policy are activated.\"},{\"question\":\"What mechanisms ensure non-interference for gating and aggregation?\",\"answer\":\"The paper designs secure gating functions and expert aggregation methods so that their decisions depend only on accessible security domains, preventing leakage from inaccessible ones.\"}]","Information Flow Control in Machine Learning through Modular Model Architecture | PDF",1785684537,45,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"information-flow-control-in-machine-learning-through-modular-model-architecture","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/information-flow-control-in-machine-learning-through-modular-model-architecture/118617/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05","2026-08-02",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What security problem does the paper address in machine learning?","Question",{"text":76,"@type":77},"It addresses how ML models can leak sensitive information from training data through inference-time outputs, especially when different users can access different subsets of training data.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the proposed IFC architecture limit information flow?",{"text":81,"@type":77},"Training data from each security domain is restricted to influence only a single expert module, and at inference time only the experts permitted by the user’s access policy are activated.",{"name":83,"@type":74,"acceptedAnswer":84},"What mechanisms ensure non-interference for gating and aggregation?",{"text":85,"@type":77},"The paper designs secure gating functions and expert aggregation methods so that their decisions depend only on accessible security domains, preventing leakage from inaccessible ones.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]