[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117625-en":3,"doc-seo-117625-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117625,962075114101,"Seraphina","https://ap-avatar.wpscdn.com/avatar/e000253a75eb197efd?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780044092746381165",8,"Research & Report","On the Robustness of Dataset Inference","Machine learning models are expensive to train and often treated as valuable intellectual property, motivating defenses against model-stealing adversaries. Ownership verification techniques such as watermarking and fingerprinting aim to prove whether a suspect model was stolen from a victim. While Dataset Inference (DI) was previously shown to be robust and efficient, this work proves DI can produce high false positives by confusing independent models trained on non-overlapping data from the same distribution, including in realistic nonlinear cases. The study also shows DI can be evaded via false negatives through adversarial training that regularizes stolen model decision boundaries, and verifies via black-box experiments that DI fails with high confidence in the hardest-to-evade setting. ","This is an electronic reprint of the original article.  \nThis reprint may differ from the original in pagination and typographic detail.  \nSzyller, Sebastian; Zhang, Rui; Liu, Jian; Asokan, N  \nOn the Robustness of Dataset Inference  \nPublished in:  \nTransactions on Machine Learning Research  \nPublished: 01/01/2023  \nDocument Version  \nPublisher's PDF, also known as Version of record  \nPublished under the following license:  \nCC BY  \nPlease cite the original version:  \nSzyller, S. , Zhang, R. , Liu, J. , & Asokan, N. (2023) . On the Robustness of Dataset Inference. Transactions on Machine Learning Research, 2023(6) . [https://openreview.net/forum?id=LKz5SqIXPJ](https://openreview.net/forum?id=LKz5SqIXPJ)  \nThis material is protected by copyright and other intellectual property rights, and duplication or sale of all or part of any of the repository collections is not permitted, except that material may be duplicated by you foryour research use or educational purposes in electronic or print form. You must obtain permission for anyother use. Electronic or print copies may not be offered, whether for sale or otherwise to anyone who is not an authorised user.  \nOn the Robustness of Dataset Inference  \nSebastian Szyller  \nAalto University  \nRui Zhang  \nZhejiang University  \nJian Liu  \nZhejiang University  \nN. Asokan  \nUniversity of Waterloo & Aalto University  \n[contact@sebszyller. com](contact@sebszyller. com)  \n[zhangrui98@zju. edu. cn](zhangrui98@zju. edu. cn)  \n[liujian2411@zju. edu. cn](liujian2411@zju. edu. cn)  \n[asokan@acm. org](asokan@acm. org)  \nReviewed on OpenReview: ht [tp s: // op en re vi ew .n et /f or um ?i](tp s: // op en re vi ew .n et /f or um ?i) d= LK z5 Sq IX PJ  \nAbstract  \nMachine learning (ML) models are costly to train as they can require a significant amount of data, computational resources and technical expertise. Thus, they constitute valuable intellectual property that needs protection from adversaries wanting to steal them. Ownership verification techniques allow the victims of model stealing attacks to demonstrate that a suspect model was in fact stolen from theirs.  \nAlthough a number of ownership verification techniques based on watermarking or fingerprinting have been proposed, most of them fall short either in terms of security guarantees (well-equipped adversaries can evade verification) or computational cost. A fingerprinting technique, Dataset Inference (DI), has been shown to offer better robustness and efficiency than prior methods.  \nThe authors of DI provided a correctness proof for linear (suspect) models. However, in a subspace of the same setting, we prove that DI suffers from high false positives (FPs)– it can incorrectly identify an independent model trained with non-overlapping data from the same distribution as stolen. We further prove that DI also triggers FPs in realistic, non-linear suspect models. We then confirm empirically that DI in the black-box setting leads to FPs, with high confidence.  \nSecond, we show that DI also suffers from false negatives (FNs)– an adversary can fool DI (at the cost of incurring some accuracy loss) by regularising a stolen model’s decision boundaries using adversarial training, thereby leading to an FN. To this end, we demonstrate that black-box DI fails to identify a model adversarially trained from a stolen dataset – the setting where DI is the hardest to evade.  \nFinally, we discuss the implications of our findings, the viability of fingerprinting-based ownership verification in general, and suggest directions for future work.  \n1 Introduction  \nMachine learning (ML) models are being developed and deployed at an increasingly faster rate and in several application domains. For many companies, they are not just a part of the technological stack that offers an edge over the competitors but a core business offering. Hence, ML models constitute valuable intellectual property that needs to be protected.  \nModel stealing is considered one of the most se","cbCairTK27aylcC6","https://ap.wps.com/l/cbCairTK27aylcC6","pdf",1243832,1,20,"English","en",105,"# Introduction\n## Model stealing and ownership verification\n## Watermarking vs. fingerprinting\n## Dataset Inference (DI) robustness and failure modes\n## False positives in linear and nonlinear settings\n## False negatives via adversarial training and black-box evasion","[{\"question\":\"What problem does the paper address in machine learning security?\",\"answer\":\"It analyzes how reliable Dataset Inference is for ownership verification against model stealing attacks, where adversaries attempt to obtain functionally equivalent copies of victim models.\"},{\"question\":\"How does the paper show Dataset Inference can fail?\",\"answer\":\"The paper proves DI can suffer from high false positives, incorrectly identifying independent models trained on non-overlapping data as stolen, and it also triggers false positives for realistic nonlinear suspect models.\"},{\"question\":\"How can an attacker cause false negatives against DI?\",\"answer\":\"An attacker can fool DI by regularizing a stolen model’s decision boundaries using adversarial training, which can introduce an accuracy cost while leading DI to miss the stolen model.\"}]","On the Robustness of Dataset Inference | PDF",1785677375,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"on-the-robustness-of-dataset-inference","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/on-the-robustness-of-dataset-inference/117625/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in machine learning security?","Question",{"text":75,"@type":76},"It analyzes how reliable Dataset Inference is for ownership verification against model stealing attacks, where adversaries attempt to obtain functionally equivalent copies of victim models.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the paper show Dataset Inference can fail?",{"text":80,"@type":76},"The paper proves DI can suffer from high false positives, incorrectly identifying independent models trained on non-overlapping data as stolen, and it also triggers false positives for realistic nonlinear suspect models.",{"name":82,"@type":73,"acceptedAnswer":83},"How can an attacker cause false negatives against DI?",{"text":84,"@type":76},"An attacker can fool DI by regularizing a stolen model’s decision boundaries using adversarial training, which can introduce an accuracy cost while leading DI to miss the stolen model.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":29,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":21,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":21,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]