[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-126117-en":3,"doc-seo-126117-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":11,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},126117,5909887254083,"Miles","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Analyzing Threats of Large-Scale Machine Learning Systems","Large-scale machine learning systems such as ChatGPT rapidly reshape how people interact with and trust digital media, but they also create a dual-use dilemma. The same capabilities can be exploited by untrustworthy actors to cause intentional harm, including unintended disclosure of private training information and generation of content that undermines authenticity. This thesis analyzes threats through two lenses: data leakage via memorization and misuse when users act outside intended purposes. Five projects assess privacy and security risks and evaluate countermeasures, including risks under differential privacy, and the robustness of fingerprinting and watermarking for provenance detection.","Analyzing Threats of Large-Scale Machine Learning Systems  \nby  \nNils Lukas  \nA thesis  \npresented to the University of Waterloo  \nin fulfillment of the  \nthesis requirement for the degree of  \nDoctor of Philosophy  \nin  \nComputer Science  \nWaterloo, Ontario, Canada, 2024  \n© Nils Lukas 2024  \nExamining Committee Membership  \nThe following served on the Examining Committee for this thesis. The decision of the Examining Committee is by majority vote.  \nExternal Examiner: Alina Oprea  \nAssociate Professor, Khoury College of Computer Sciences Northeastern University  \nSupervisor(s): Florian Kerschbaum  \nProfessor, David R. Cheriton School of Computer Science University of Waterloo  \nInternal Member: N. Asokan  \nProfessor, David R. Cheriton School of Computer Science University of Waterloo  \nYaoliang Yu  \nAssociate Professor, David R. Cheriton School of Computer Science University of Waterloo  \nInternal-External Member: Arie Gurfinkel  \nProfessor, Department of Electrical and Computer Engineering University of Waterloo  \nAuthor’s Declaration  \nThis thesis consists of material, all of which I authored or co-authored: see the Statement of Contributions included in the thesis. This is a true copy of the thesis, including any required final revisions, as accepted by my examiners.  \nI understand that my thesis may be made electronically available to the public.  \nStatement of Contributions  \nThis dissertation includes first-authored and peer-reviewed work that has appeared in conference proceedings published by the Institute of Electrical and Electronics Engineers (IEEE), the Advanced Computing Systems Association (USENIX), and the International Conference on Learning Representations (ICLR) .  \nThe following is a list that serves as a declaration of the works included in this dissertation. I expand and revise material from the original publications.  \nPortions of Chapter 3  \nNils Lukas, Ahmed Salem, Robert Sim, Shruti Tople, Lukas Wutschitz, and Santiago Zanella-B´eguelin. 2023. Analyzing Leakage of Personally Identifiable Information in Language Models. Published in the 44th Proceedings of the 2023 IEEE Symposium on Security and Privacy.  \nPortions of Chapter 4  \nNils Lukas, Edward Jiang, Xinda Li and Florian Kerschbaum. 2022. SoK: How Robust is Deep Neural Network Image Classification Watermarking? Published in the 43rd Proceedings of the 2022 IEEE Symposium on Security and Privacy.  \nPortions of Chapter 5  \nNils Lukas, Yuxuan Zhang and Florian Kerschbaum. 2021. Deep Neural Network Fingerprinting with Conferrable Adversarial Examples. Published in the 9th Proceedings of the 2021 International Conference on Learning Representations.  \nPortions of Chapter 6  \nNils Lukas and Florian Kerschbaum. 2023. PTW: Pivotal Tuning Watermarking for PreTrained Image Generators. Published in the 32nd USENIX Security Symposium.  \nPortions of Chapter 7  \nNils Lukas, Abdulrahman Diaa, Lucas Fenaux and Florian Kerschbaum. 2023. Leveraging Optimization for Adaptive Attacks on Image Watermarks. Published in the 12th Proceedings of the 2024 International Conference on Learning Representations.  \nAbstract  \nLarge-scale machine learning systems such as ChatGPT rapidly transform how we interact with and trust digital media. However, the emergence of such a powerful technology faces a dual-use dilemma. While it can have many positive societal impacts in providing equitable access to information, ML systems can also be misused by untrustworthy entities to cause intentional harm. For example, a system could unintentionally disclose private information about its training data and jeopardize the privacy of individuals in the training data. The system’s generated content could also be misused for unethical purposes, such as eroding trust in digital media by misrepresenting generating content as authentic. Providing untrustworthy users with these new capabilities could amplify potential negative consequences emerging through this technology, such as a proliferation o","cbCaifQZPQshemFq","https://ap.wps.com/l/cbCaifQZPQshemFq","pdf",10300844,1,207,"English","en",105,"# Abstract\n# Introduction and Dual-Use Threats\n## Data Leakage and Memorization\n## Misuse and Unintended Use\n# Thesis Contributions and Project Overview\n## Privacy Risks and Differential Privacy\n## Provenance Detection via Fingerprinting and Watermarking\n# Contributions by Chapter (Portions 3–7)","[{\"question\":\"What dual-use dilemma does the thesis focus on?\",\"answer\":\"It examines how large-scale ML systems can deliver societal benefits while also enabling intentional harm through misuse and privacy/security failures.\"},{\"question\":\"How does the thesis categorize the analyzed threats?\",\"answer\":\"It analyzes threats from two perspectives: data leakage when models memorize private information during training, and misuse when users do not use the system as intended.\"},{\"question\":\"Which countermeasures does the thesis evaluate?\",\"answer\":\"It assesses privacy risks related to extracting personally identifiable information under differential privacy, and evaluates the effectiveness and robustness of fingerprinting and watermarking methods for provenance detection.\"}]","Analyzing Threats of Large-Scale Machine Learning Systems | PDF",1785903254,522,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"analyzing-threats-of-large-scale-machine-learning-systems","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/analyzing-threats-of-large-scale-machine-learning-systems/126117/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-23","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":11},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What dual-use dilemma does the thesis focus on?","Question",{"text":76,"@type":77},"It examines how large-scale ML systems can deliver societal benefits while also enabling intentional harm through misuse and privacy/security failures.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the thesis categorize the analyzed threats?",{"text":81,"@type":77},"It analyzes threats from two perspectives: data leakage when models memorize private information during training, and misuse when users do not use the system as intended.",{"name":83,"@type":74,"acceptedAnswer":84},"Which countermeasures does the thesis evaluate?",{"text":85,"@type":77},"It assesses privacy risks related to extracting personally identifiable information under differential privacy, and evaluates the effectiveness and robustness of fingerprinting and watermarking methods for provenance detection.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]