[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85962-en":3,"doc-seo-85962-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85962,13056703019404,"Miles","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Observation-Level Watermarking and Detection for Tabular Data","Generative AI increases the need for reliable authenticity verification, data ownership protection, and proper attribution. While watermarking is mature for images and text, watermarking tabular data is comparatively under-explored, especially for discrete, categorical, and mixed-variable settings. This work introduces STAMP (Single-observation Tabular Attribution and Marking Procedure) to insert and preserve diverse distributions, provides a detection mechanism that works with a single sample, and proves asymptotic consistency and accurate detection. Simulations and real-data applications confirm robustness under subsetting while maintaining fidelity.","arXiv :2607 . 10554v1 [ stat .ME] 12 Jul 2026  \nObservation-Level Watermarking and Detection for  \nTabular Data  \nDongyu Cui∗ Xuan Bi†  \nAbstract  \nWith the development of generative AI, watermarking techniques have been widely used to detect the authenticity of AI-generated data and protect the rights of users and creators. While it is already well applied in data types including imaging and text data, watermarking tabular data is still under-explored. Existing methods primarily focus on numerical data, leaving discrete, categorical, and mixed data less studied. In this work, we propose STAMP (Single-observation Tabular Attribution and Marking Procedure), a novel framework for watermarking tabular data that can accommodate and preserve a wide range of distributions. We also develop a corresponding detection mechanism, which can reliably identify watermarks even when the sample size is as small as one. We establish theoretical guarantees for asymptotic consistency and detection accuracy. Finally, through extensive simulation studies and two real-data applications, we demonstrate that the proposed method is effective and robust to subsetting, while maintaining data fidelity and a high detection rate.  \n1 Introduction  \nIn modern society, data are produced and shared at an unprecedented rate. With the rapid development of generative AI techniques, it has become increasingly easy to generate various types of data,  \n∗ School of Statistics, University of Minnesota  \n†Carlson School of Management, University of Minnesota, email: [xbi@umn.edu](xbi@umn.edu)  \nincluding images, text, and tabular data. This development has raised several important issues. The first issue is about data authenticity. On the one hand, AI-generated data can be highly realistic and easily mistaken for real data. On the other hand, datasets may be provided without a clear source and misrepresented as originating from reputable individuals or institutions. Both cases highlight the need for reliable authenticity verification. Second, data may be used by others without the creator’s permission, raising concerns about ownership and unauthorized use. Third, even when data are shared, creators may want to seek proper attribution, as malicious recipients may redistribute the data while falsely claiming it as their own. Last, ensuring traceability is also important, so that the origin of a dataset can be identified and responsible parties can be held accountable for misuse or leakage.These needs call for techniques to make data identifiable, and thereby protect the rights of users and creators.  \nWatermarking is a widely used technique to address these issues. It involves embedding a unique identifier into the data, which later can be used to verify the data’s origin. Specifically, a watermarking method usually consists of two steps: watermark insertion and watermark detection. First, a unique identifier (i.e., a key) is inserted into the data to make it identifiable. Second, a detection procedure is used to determine the presence of the key.  \nIn recent years, watermarking methods have been proposed for different data types and modalities. For image data, commonly used methods include frequency-domain methods (e.g., signal encryption after Fourier transformation) [2, 10] and recent neural-network-based methods [7, 20, 23, 29] . For text data, watermarking techniques include synonym substitution [24], sentence reordering [3], the green-red list [12], or character-level modifications [19] . Specifically, for watermarking text generated by large language models, several methods have been proposed, such as modifying the token selection process during text generation or introducing specific patterns in the generated text [1, 11, 13] . These methods aim to make the watermark robust against paraphrasing and other text transformations while preserving the readability and coherence of the text. On the other hand, designing and detecting watermarks may encounter several chal","cbCaiuzCG5NlVYqC","https://ap.wps.com/l/cbCaiuzCG5NlVYqC","pdf",510341,1,53,"English","en",105,"# Introduction\n## Watermarking background and problem motivation\n## Prior work across imaging and text\n## Existing tabular watermarking methods\n## Key challenges and contributions","[{\"question\":\"What problem does this work address in the context of generative AI?\",\"answer\":\"It addresses the need for authenticity verification, ownership protection, attribution, and traceability for data produced and shared via generative AI, especially when tabular data may be misrepresented or reused without permission.\"},{\"question\":\"What is STAMP and what does it enable for tabular data watermarking?\",\"answer\":\"STAMP (Single-observation Tabular Attribution and Marking Procedure) is proposed as a unified watermarking framework for tabular data that can accommodate and preserve a wide range of distributions, including discrete, categorical, and mixed data types.\"},{\"question\":\"How does the proposed detection method perform under limited sample sizes?\",\"answer\":\"The work develops a corresponding detection mechanism that can reliably identify watermarks even when the sample size is as small as one, supported by theoretical guarantees for asymptotic consistency and detection accuracy.\"}]",1784207405,134,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"observation-level-watermarking-and-detection-for-tabular-data","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/observation-level-watermarking-and-detection-for-tabular-data/85962/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does this work address in the context of generative AI?","Question",{"text":75,"@type":76},"It addresses the need for authenticity verification, ownership protection, attribution, and traceability for data produced and shared via generative AI, especially when tabular data may be misrepresented or reused without permission.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is STAMP and what does it enable for tabular data watermarking?",{"text":80,"@type":76},"STAMP (Single-observation Tabular Attribution and Marking Procedure) is proposed as a unified watermarking framework for tabular data that can accommodate and preserve a wide range of distributions, including discrete, categorical, and mixed data types.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the proposed detection method perform under limited sample sizes?",{"text":84,"@type":76},"The work develops a corresponding detection mechanism that can reliably identify watermarks even when the sample size is as small as one, supported by theoretical guarantees for asymptotic consistency and detection accuracy.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]