[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85248-en":3,"doc-seo-85248-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85248,1374391974564,"Clementine","https://ap-avatar.wpscdn.com/avatar/14000253aa45c000a9e?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779874745381141002",8,"Research & Report","Can Watermarking Techniques Help Prevent LLM Model Stealing","Model stealing attacks enable extraction of precise information from black-box commercial language models through API access. The work proposes defense methods targeting a recent Carlini et al. 2024b style attack and extensions aimed at revealing hidden-layer dimensions. Inspired by watermarking, the defenses perturb the models’ logits using structured, targeted perturbations derived from private hashes. Experiments show the approach is effective against rank-dimension extraction while limiting model quality degradation across configurations.","Can Watermarking Techniques Help Prevent LLM Model Stealing?  \nElette Boyle 1 , MohammadTaghi Hajiaghayi2 , Keivan Rezaei2 , Suho Shin2 and Amos Stern 1  \n1Reichman University  \n2University of Maryland  \narXiv :2607 . 10794v 1 [ cs .CR] 12 Jul 2026  \nAbstract  \nModel stealing attacks have recently been introduced, enabling the extraction of precise information from black-box commercial language models.  \nIn this work, we propose defense methods against a recent attack of [Carlini et al., 2024b] and extensions for extracting the hidden layer dimension of production language models. Our methods are inspired by watermarking techniques that perturb the logits layer of these models to prevent such attacks.  \nWe provide empirical experiments demonstrating the effectiveness of the proposed defense versus model quality degradation across various configurations, and propose an effective defense against such attacks while preserving model utility.  \n1 Introduction  \nCommercial language models (LMs) such as Gemini, GPT- 4, and Claude [Achiam et al., 2023; Team et al., 2023; Anthropic, 2024], which are publicly accessible, require substantial time and resources for training. Consequently, details about these models, including training data, methodologies, architectures, and parameters, are typically not disclosed. However, recent reports highlight the risks of model stealing, where unauthorized actors attempt to extract proprietary knowledge. For example, recent article [Reuters, 2025] indicate the possibility of knowledge being extracted from OpenAI’s API output to produce a competing product.  \nThese models operate through API access, allowing users to input prompts and receive predictions. [Carlini et al., 2024b] proposed techniques to extract precise information from various black-box LMs, such as their hidden state dimensions and final projection layers. This raises serious concerns that extension of these attacks could potentially reveal even more sensitive information.  \nSpecifically, [Carlini et al., 2024b] present a scalable approach targeting production LMs. Unlike earlier model stealing methods [Carlini et al., 2020; Carlini et al., 2024a; Rolnick and Kording, 2020], it does not depend on specific activation functions in the model or network structures, making it both versatile and a significant concern. This highlights the need for robust defenses against such attacks. The method described in [Carlini et al., 2024b] leverages a model’s API  \nFigure 1: Our framework of perturbing logits to prevent dimensionextraction model stealing attacks.  \nto extract logit vectors of next-token-prediction on a set of prompts. These logits are then used to construct a matrix, whose properties are used for extracting hidden dimension, and later the last projection layer in a black-box LMs. While the specific API queries that enable logit vector reconstruction in the [Carlini et al., 2024b] attack have since been removed from major models, 1 their attack indicates a clear danger and potential vulnerability. In particular, the hidden latent dimension of the target model is a critical component for [Carlini et al., 2024b], as it serves as the foundation for stealing the model’s last layer.  \nIn this work, we propose a structured perturbation of logits to counter rank-estimation attacks, inspired by watermarking techniques [Kirchenbauer et al., 2023] . Unlike prior methods that sample and add independent noise, we seed targeted perturbations using a private hash of the model’s hidden latents. Our approach significantly increases the complexity of attacks attempting to break the defense, while preserving the model’s quality. Our overall framework is given in Figure 1 .  \nWe remark that watermarking has been used for post-hoc detection of behavioral cloning [Zhao et al., 2023] . Our method takes a very different approach, proactively disrupting rank-based extraction attacks of architectural details such as hidden dimension. (See Related Work in Appe","cbCaiftuey0fOUj2","https://ap.wps.com/l/cbCaiftuey0fOUj2","pdf",16903843,3,1,16,"English","en",105,"# Introduction\n## Summary of results\n# Attack methods\n## PCA\n## PCA with averaging\n## Robust PCA\n# Defense mechanisms","[{\"question\":\"What problem does the paper address?\",\"answer\":\"The paper addresses model stealing attacks that extract sensitive information from black-box commercial language models via API queries, including attempts to recover hidden-layer dimensions.\"},{\"question\":\"How do the proposed defenses work?\",\"answer\":\"The defenses use watermarking-inspired, structured perturbations of the logits layer so that rank-estimation attacks become significantly harder to break while keeping the model useful.\"},{\"question\":\"What evidence is provided for effectiveness?\",\"answer\":\"Empirical experiments demonstrate improved defense effectiveness against rank/dimension extraction with reduced model quality degradation across multiple configurations.\"}]",1784202067,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"can-watermarking-techniques-help-prevent-llm-model-stealing","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/can-watermarking-techniques-help-prevent-llm-model-stealing/85248/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address?","Question",{"text":75,"@type":76},"The paper addresses model stealing attacks that extract sensitive information from black-box commercial language models via API queries, including attempts to recover hidden-layer dimensions.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do the proposed defenses work?",{"text":80,"@type":76},"The defenses use watermarking-inspired, structured perturbations of the logits layer so that rank-estimation attacks become significantly harder to break while keeping the model useful.",{"name":82,"@type":73,"acceptedAnswer":83},"What evidence is provided for effectiveness?",{"text":84,"@type":76},"Empirical experiments demonstrate improved defense effectiveness against rank/dimension extraction with reduced model quality degradation across multiple configurations.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":29,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]