[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-126305-en":3,"doc-seo-126305-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":11,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},126305,2336475104957,"Seraphina","https://ap-avatar.wpscdn.com/avatar/22000c4c6bd8a5076e1?x-image-process=image/resize,m_fixed,w_180,h_180&k=1786593998035447633",8,"Research & Report","SoK: Dataset Copyright Auditing in Machine Learning Systems - Research article","As machine learning systems become widely deployed and larger models drive massive data demand, dataset copyright infringement and misuse risks increase, from unauthorized online artworks to facial images used for training. Existing auditing approaches differ in assumptions and capabilities, limiting meaningful comparison, while robustness evaluations often cover only parts of the ML pipeline and fail to reflect real-world deployment performance. This work proposes a practical deployment perspective, organizes methods into intrusive and non-intrusive auditing, compares watermark and fingerprint strategies, and synthesizes reference tables and unresolved research directions.","VU Research Portal  \nSoK: Dataset Copyright Auditing in Machine Learning Systems  \nDu, Linkang; Zhou, Xuanru; Chen, Min; Zhang, Chusong; Su, Zhou; Cheng, Peng; Chen, Jiming; Zhang, Zhikun  \npublished in  \n2025 IEEE Symposium on Security and Privacy (SP)  \n2025  \nDOI (link to publisher)  \n10.1109/SP61157.2025.00025  \ndocument version  \nPublisher's PDF, also known as Version of record  \ndocument license  \nArticle 25fa Dutch Copyright Act  \nLink to publication in VU Research Portal  \ncitation for published version (APA)  \nDu, L. , Zhou, X. , Chen, M. , Zhang, C. , Su, Z. , Cheng, P. , Chen, J. , & Zhang, Z. (2025) . SoK: Dataset Copyright Auditing in Machine Learning Systems. In M. Blanton, W. Enck, & C. Nita-Rotaru (Eds. ), 2025 IEEE Symposium on Security and Privacy (SP): [Proceedings](pp. 2076-2094) . (Proceedings-IEEE Symposium on Security and Privacy; Vol. 2025) . Institute of Electrical and Electronics Engineers Inc..  \n[https://doi.org/10.1109/SP61157.2025.00025](https://doi.org/10.1109/SP61157.2025.00025)  \nGeneral rights  \nCopyright and moral rights for the publications made accessible in the public portal are retained by the authors and/or other copyright owners and it is a condition of accessing publications that users recognise and abide by the legal requirements associated with these rights.  \n• Users may download and print one copy of any publication from the public portal for the purpose of private study or research.  \n• You may not further distribute the material or use it for any profit-making activity or commercial gain  \n• You may freely distribute the URL identifying the publication in the public portal  \nTake down policy  \nIf you believe that this document breaches copyright please contact us providing details, and we will remove access to the work immediately and investigate your claim.  \nE-mail address:  \n[vuresearchportal.ub@vu.nl](vuresearchportal.ub@vu.nl)  \n[Download date: 16](Download date: 16) . May. 2026  \n2025 IEEE Symposium on Security and Privacy (SP) ©2025 IEEE DOI: 10.1109/SP61157.2025.00025| 979-8-3315-2236-0/25/$31.00 |   \n2025 IEEE Symposium on Security and Privacy (SP)  \nSoK: Dataset Copyright Auditing in Machine Learning Systems  \nLinkang Du 1* Xuanru Zhou2* Min Chen3 Chusong Zhang2 Zhou Su 1 Peng Cheng2 Jiming Chen2 ,4 Zhikun Zhang2\\#  \n1 Xi’an Jiaotong University 2Zhejiang University 3 Vrije Universiteit Amsterdam 4 Hangzhou Dianzi University  \nAbstract—As the implementation of machine learning (ML) systems becomes more widespread, especially with the introduction of larger ML models, we perceive a spring demand for massive data. However, it inevitably causes infringement and misuse problems with the data, such as using unauthorized online artworks or face images to train ML models. To address this problem, many efforts have been made to audit the copyright of the model training dataset. However, existing solutions vary in auditing assumptions and capabilities, making it difficult to compare their strengths and weaknesses. In addition, robustness evaluations usually consider only part of the ML pipeline and hardly reflect the performance of algorithms in real-world ML applications. Thus, it is essential to take a practical deployment perspective on the current dataset copyright auditing tools, examining their effectiveness and limitations. Concretely, we categorize dataset copyright auditing research into two prominent strands: intrusive methods and non-intrusive methods, depending on whether they require modifications to the original dataset. Then, we break down the intrusive methods into different watermark injection options and examine the non-intrusive methods using various fingerprints. To summarize our results, we offer detailed reference tables, highlight key points, and pinpoint unresolved issues in the current literature. By combining the pipeline in ML systems and analyzing previous studies, we highlight several future directions to make auditing tools more suitable for realworl","cbCaigYrWWAWnNzW","https://ap.wps.com/l/cbCaigYrWWAWnNzW","pdf",1213895,1,20,"English","en",105,"# Abstract\n# Introduction\n# Dataset Copyright Auditing: Intrusive Methods\n## Watermark Injection Options\n# Dataset Copyright Auditing: Non-intrusive Methods\n## Fingerprint-Based Approaches\n# Synthesis and Open Issues\n## Reference Tables and Future Directions","[{\"question\":\"Why is dataset copyright auditing increasingly important for ML systems?\",\"answer\":\"ML systems need large-scale datasets, which increases the chance of copyright infringement and misuse. Large models further amplify demand for massive data, raising risks such as using unauthorized images for training.\"},{\"question\":\"How does the paper categorize dataset copyright auditing research?\",\"answer\":\"It divides auditing into two prominent strands: intrusive methods and non-intrusive methods. The distinction depends on whether they require modifications to the original dataset.\"},{\"question\":\"What limitations do existing evaluations have, according to the paper?\",\"answer\":\"Robustness evaluations often consider only part of the ML pipeline, making it hard to reflect algorithm performance in real-world ML applications. This limits assessment of effectiveness and constraints for deployment.\"}]","SoK: Dataset Copyright Auditing in Machine Learning Systems - Research article | PDF",1785904362,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"sok-dataset-copyright-auditing-in-machine-learning-systems-research-article","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/sok-dataset-copyright-auditing-in-machine-learning-systems-research-article/126305/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-22","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":11},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why is dataset copyright auditing increasingly important for ML systems?","Question",{"text":76,"@type":77},"ML systems need large-scale datasets, which increases the chance of copyright infringement and misuse. Large models further amplify demand for massive data, raising risks such as using unauthorized images for training.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the paper categorize dataset copyright auditing research?",{"text":81,"@type":77},"It divides auditing into two prominent strands: intrusive methods and non-intrusive methods. The distinction depends on whether they require modifications to the original dataset.",{"name":83,"@type":74,"acceptedAnswer":84},"What limitations do existing evaluations have, according to the paper?",{"text":85,"@type":77},"Robustness evaluations often consider only part of the ML pipeline, making it hard to reflect algorithm performance in real-world ML applications. This limits assessment of effectiveness and constraints for deployment.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,115,120,123,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":29,"slug":114},6,"Technology","technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":21,"slug":126},9,"Religion & Spirituality","religion-spirituality",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":21,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":107,"slug":137},19,"General","general"]