[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84550-en":3,"doc-seo-84550-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84550,34359740700684,"Finn","https://ap-avatar.wpscdn.com/avatar/1f400023980c374ae676?_k=1777273430885731487",8,"Research & Report","Watermarking for Proprietary Dataset Protection","A growing body of work argues that training data membership inference is fundamentally difficult in modern language modeling. This study contends that output watermarking provides the right mechanism to make generative-model training membership testing more tractable, leveraging evidence that language models exhibit residual watermark “radioactivity” when trained on partially watermarked datasets. A watermark-based dataset inference method is compared against loss-based membership inference and can reach comparable membership detection performance under sufficiently high subset exposure under alternate assumptions.","Watermarking for Proprietary Dataset Protection  \nJohn Kirchenbauer 1 Brian R. Bartoldson 2 Bhavya Kailkhura 2 Tom Goldstein 1  \narXiv :2607 .00325v 1 [ cs .LG] 1 Jul 2026  \nAbstract  \nA growing body of literature suggests that training data membership inference problems are fundamentally hard tasks in modern language modeling settings. We argue that output watermarking techniques are the right gadget to make training membership tests for generative models more tractable, based on prior results showing that language models exhibit residual watermark “radioactivity” under partially watermarked training datasets. We pit a watermark-based dataset inference approach head-to-head against traditional loss-based membership inference methods and show that watermarking can achieve comparable membership detection performance when subset exposure is high enough, under an alternate set of assumptions.  \n1. Introduction  \nModern language models perform complex knowledge work of growing economic value, but the regulatory frameworks governing fair use of the web-scraped data they train on remain underdeveloped. Recent litigation suggests content owners like news websites and independent authors maybe entitled to compensation for inclusion of their datasets in large-scale AI model training. Answering such datause questions in high-stakes settings requires a concrete definition of what it means to test whether some data was included in a model’s training dataset.  \nWhile the fully general training data attribution problem asks how a model’s test-time behaviors are caused by specific training instances, the question at hand in contemporary fair-use deliberations is actually just membership. Membership inference attacks (MIAs) ask whether a specific sample was in a model’s training dataset; dataset inference attacks (DIAs) generalize this to whole collections. As the more relevant setting for IP and generative-model training disputes, in this work we study the DIA setup but refer to both problem types collectively as membership problems.  \n1University of Maryland 2Lawrence Livermore National Labs. Correspondence to: John Kirchenbauer \u003C[jkirchen@umd.edu](jkirchen@umd.edu) >.  \nPresented at the ICML 2026 Workshop on Trustworthy AI4GOOD, Seoul, South Korea. PMLR 306, 2026 .  \nFigure 1. To protect a proprietary dataset from unauthorized use in training, the dataset owner (attacker) paraphrases their documents with a secret watermark key. To perform dataset inference, the attacker tests the suspect model’s predictions for evidence of the watermark key. The watermark detection test is used to conclude whether their protected data was included in the training dataset.  \nWhat makes training dataset membership tests challenging in modern settings? Performing membership tests on modern generative models is harder than it was in discriminative settings for classifiers which mapped inputs to as few as 10 or 100 classes (i.e. the setting of Shokri et al. (2017)) . The variable-length output space matches the input cardinality, making analysis complicated, and the billion-parameter sizes of modern models make classical attribution tools like influence functions difficult to apply (Koh & Liang, 2017 ; Ilyas et al., 2022 ; Park et al., 2023) . Even the definition of “a sample” is arbitrary at pretraining scale (documents, sentences, words), and individual samples overlap heavily in terms of shared n-grams, which destabilizes existing membership tests (Duan et al., 2024) . Modern generative models also memorize and generalize along both semantic and stylistic lines, blurring the line between membership and generation attributability: sample membership is often neither necessary nor sufficient for a target output to be produced (Liu et al., 2025), and memorization rates vary widely (Cooper et al., 2025) .  \nProactive Interventions for Reliable Dataset Inference.  \nOur work attempts to circumvent some of the difficulties described above by proactively but sparingl","cbCailKQ3F20CPm1","https://ap.wps.com/l/cbCailKQ3F20CPm1","pdf",2306807,1,32,"English","en",105,"# Introduction\n## Membership and dataset inference problems\n# Proactive interventions for reliable dataset inference\n## Output watermarking layered on paraphrasers\n# Relationship to prior work\n## Prior folding and detector construction","[{\"question\":\"What problem does the paper address?\",\"answer\":\"The paper addresses training-data dataset inference, i.e., determining whether a collection of data was included in a generative model’s training set, framed alongside membership inference.\"},{\"question\":\"How does output watermarking help make membership tests easier?\",\"answer\":\"It uses watermarked decoding layered on a paraphraser so candidate samples carry a detectable watermark signature; if the target model was trained on these samples, the model reproduces a shadow of the key-specific signature at test time.\"},{\"question\":\"How is the watermark-based approach evaluated against traditional methods?\",\"answer\":\"The study compares a watermark-based dataset inference approach with loss-based membership inference methods and reports comparable membership detection performance when marked subset exposure is high enough under alternate assumptions.\"}]",1784196615,81,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"watermarking-for-proprietary-dataset-protection","",{"@graph":35,"@context":84},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/watermarking-for-proprietary-dataset-protection/84550/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What problem does the paper address?","Question",{"text":74,"@type":75},"The paper addresses training-data dataset inference, i.e., determining whether a collection of data was included in a generative model’s training set, framed alongside membership inference.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does output watermarking help make membership tests easier?",{"text":79,"@type":75},"It uses watermarked decoding layered on a paraphraser so candidate samples carry a detectable watermark signature; if the target model was trained on these samples, the model reproduces a shadow of the key-specific signature at test time.",{"name":81,"@type":72,"acceptedAnswer":82},"How is the watermark-based approach evaluated against traditional methods?",{"text":83,"@type":75},"The study compares a watermark-based dataset inference approach with loss-based membership inference methods and reports comparable membership detection performance when marked subset exposure is high enough under alternate assumptions.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":105,"slug":137},19,"General","general"]