[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"detail-sidebar-cat-1-en-105":3,"doc-seo-275863-105":53,"doc-detail-275863-en":126},{"code":4,"msg":5,"data":6},0,"success",[7,14,19,24,29,34,39,44,49],{"id":8,"doc_module":9,"doc_module_name":10,"category_name":11,"show_sort_weight":12,"slug":13},11,1,"Template","Presentations",90,"presentations",{"id":15,"doc_module":9,"doc_module_name":10,"category_name":16,"show_sort_weight":17,"slug":18},12,"Resumes",80,"resumes",{"id":20,"doc_module":9,"doc_module_name":10,"category_name":21,"show_sort_weight":22,"slug":23},14,"Invoices",70,"invoices",{"id":25,"doc_module":9,"doc_module_name":10,"category_name":26,"show_sort_weight":27,"slug":28},15,"Posters",60,"posters",{"id":30,"doc_module":9,"doc_module_name":10,"category_name":31,"show_sort_weight":32,"slug":33},16,"Social Media",50,"social-media",{"id":35,"doc_module":9,"doc_module_name":10,"category_name":36,"show_sort_weight":37,"slug":38},17,"Forms",40,"forms",{"id":40,"doc_module":9,"doc_module_name":10,"category_name":41,"show_sort_weight":42,"slug":43},18,"Letters",30,"letters",{"id":45,"doc_module":9,"doc_module_name":10,"category_name":46,"show_sort_weight":47,"slug":48},21,"Paper Templates",5,"papers-templates",{"id":50,"doc_module":9,"doc_module_name":10,"category_name":51,"show_sort_weight":4,"slug":52},158,"General","general-158",{"code":4,"msg":54,"data":55},"ok",{"site_id":56,"language":57,"slug":58,"title":59,"keywords":60,"description":61,"schema_data":62,"social_meta":119,"head_meta":121,"extra_data":123,"updated_unix":125},105,"en","alpaca-against-vicuna-using-llms-to-uncover-memorization-of-llms","ALPACA AGAINST VICUNA - Using LLMs to Uncover Memorization of LLMs","","The paper investigates how instruction-tuning influences memorization and the ability to discover pre-training data in large language models (LLMs). It introduces a black-box prompt optimization strategy where an attacker LLM agent iteratively constructs instruction-based prompts via rejection sampling to reduce direct overlap with training data while maximizing overlap between a victim model’s outputs and training continuations. The approach yields 23.7% more overlap than strong baselines and analyzes analytical and classifier-based attack settings. Results show instruction-tuned models may leak training data as much as or more than base models, with leakage extendable beyond original contexts.",{"@graph":63,"@context":118},[64,80,101],{"@type":65,"itemListElement":66},"BreadcrumbList",[67,71,74,77],{"item":68,"name":69,"@type":70,"position":9},"https://docshare.wps.com","Home","ListItem",{"item":72,"name":10,"@type":70,"position":73},"https://docshare.wps.com/template/",2,{"item":75,"name":51,"@type":70,"position":76},"https://docshare.wps.com/template/general/",3,{"item":78,"name":59,"@type":70,"position":79},"https://docshare.wps.com/template/alpaca-against-vicuna-using-llms-to-uncover-memorization-of-llms/275863/",4,{"url":78,"name":59,"@type":81,"image":82,"author":87,"headline":59,"publisher":90,"fileFormat":93,"inLanguage":57,"description":61,"dateModified":94,"datePublished":95,"encodingFormat":93,"isAccessibleForFree":96,"interactionStatistic":97},"DigitalDocument",{"url":83,"@type":84,"width":85,"height":86},"https://docshare.wps.com/thumbnails/alpaca-against-vicuna-using-llms-to-uncover-memorization-of-llms/275863.png","ImageObject",442,249,{"name":88,"@type":89},"Skyler","Person",{"url":68,"name":91,"@type":92},"DocShare","Organization","application/pdf","2026-09-20","2026-09-15",true,{"@type":98,"interactionType":99,"userInteractionCount":9},"InteractionCounter",{"@type":100},"ViewAction",{"@type":102,"mainEntity":103},"FAQPage",[104,110,114],{"name":105,"@type":106,"acceptedAnswer":107},"What impact does instruction-tuning have on memorization in LLMs, according to the paper?","Question",{"text":108,"@type":109},"Instruction-tuning can increase memorization and the discoverability of pre-training data in aligned models, enabling instruction-based prompts to uncover higher memorization levels than direct prompting with training data.","Answer",{"name":111,"@type":106,"acceptedAnswer":112},"How does the proposed black-box prompt optimization method work?",{"text":113,"@type":109},"An attacker LLM agent uses an iterative rejection-sampling loop to craft instruction-based prompts, minimizing prompt overlap with training data while maximizing overlap between the victim model’s outputs and training continuations via a reward function.",{"name":115,"@type":106,"acceptedAnswer":116},"What are the two attack settings evaluated in the study?",{"text":117,"@type":109},"The paper evaluates an analytical setting to estimate an empirical upper bound, with and without access to responses for prompt initialization, and a practical classifier-based method to assess memorization without access to memorized data.","https://schema.org",{"og:url":78,"og:type":120,"og:title":59,"og:site_name":91,"og:description":61},"article",{"robots":122,"canonical":78},"index,follow",{"doc_id":124,"site_id":56},275863,1789489914,{"code":4,"msg":5,"data":127},{"doc_id":124,"user_id":128,"nickname":88,"user_avatar":129,"doc_module":9,"category_id":50,"category_name":51,"doc_title":59,"doc_description":61,"doc_content":130,"file_id":131,"file_url":132,"file_type":133,"file_size":134,"view_count":9,"is_deleted":4,"is_public":9,"is_downloadable":9,"audit_status":9,"page_count":135,"language":136,"language_code":57,"site_id":56,"html_lang":57,"table_of_contents":137,"faqs":138,"seo_title":139,"seo_description":61,"update_tm":125,"read_time":140},2336464648746,"https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c","ALPACA AGAINST VICUNA: Using LLMs to Uncover Memorization of LLMs  \nAly M. Kassem1 * Omar Mahmoud2∗ Niloofar Mireshghallah3∗ Hyunwoo Kim4 Yulia Tsvetkov3 Yejin Choi5 Sherif Saad1 Santu Rana2  \n1University of Windsor 2A2I2, Deakin University  \n3University of Washington 4NVIDIA 5 Stanford University [kassem6@uwindsor.ca](kassem6@uwindsor.ca) , [o.mahmoud@deakin.edu.au](o.mahmoud@deakin.edu.au) , [niloofar@cs.washington.edu](niloofar@cs.washington.edu)  \nAbstract  \nIn this paper, we investigate the overlooked impact of instruction-tuning on memorization in large language models (LLMs), which has largely been studied in base, pre-trained models. We propose a black-box prompt optimization method where an attacker LLM agent uncovers higher levels of memorization in a victim agent, surpassing traditional approaches that prompt the model directly with training data. Using an iterative rejection-sampling process, we design instruction-based prompts that minimize overlap with training data to avoid providing direct solutions while maximizing overlap between the victim’s output and the training data to induce memorization. Our method shows 23. 7% more overlap with training data compared to state-of-the-art baselines. We explore two attack settings: an analytical approach that determines the empirical upper bound of the attack, both with and without access to responses for prompt initialization, and a practical classifierbased method for assessing memorization without access to memorized data. Our findings reveal that instruction-tuned models can expose pre-training data as much as, or more than, base models; contexts beyond the original training data can lead to leakage; and instructions generated by other LLMs open new avenues for automated attacks, which we believe require further exploration.1  \n1 Introduction  \nPre-trained language models are commonly instruction-tuned for user-facing applications to generate high-quality responses to task-oriented prompts (Ouyang et al., 2022 ; Taori et al., 2023a ; Chowdhery et al., 2023) . While extensive prior work has investigated memorization in pre-trained base LLMs and its implications for privacy, copyright, and fairness (Carlini et al., 2022 ; Bider-  \n*Equal Contribution  \n1§ [https://github.com/Alymostafa/](https://github.com/Alymostafa/)[ ](https://github.com/Alymostafa/)Instruction_based_attack  \nman et al., 2023a ; Shi et al., 2023 ; Mireshghallah et al., 2022), there is limited understanding of how instruction-tuning affects the memorization and discoverability of pre-training data in aligned models. Studies have shown that aligned LLMs can emit training data up to 150× more often than in regular operation (Nasr et al., 2023) . To address this gap, we pose the question: Can we use instruction-based prompts to uncover higher levels of memorization in aligned models? The established method of quantifying memorization (Carlini et al., 2023) assumes that a sequence d is memorized if prompting the model with the original prefix from the training data yields sequence d (ora similar sequence for approximate memorization; Biderman et al. 2023a) . However, recent findings suggest that prompts other than the original training data may trigger even higher levels of regurgitation (Schwarzschild et al., 2024) . To explore this, we propose a new optimization method, illustrated in Figure Figure 1, where an aligned language model acts as an ‘attacker,’ generating prompts that induce a victim (target) model to produce outputs more faithful to the training data. The attacker refines prompts through a feedback loop guided by a reward function that increases the overlap between the victim’s output and the ground truth. This approach is inspired by adversarial methods in computer security literature (Wang et al., 2023a) and has been effective in jailbreaking attacks (Mehrotra et al., 2023a ; Zeng et al., 2024 ; Ramesh et al., 2024) .  \nTo evaluate our approach, we draw parallels between safety jailbreaki","cbCais2F0YAC7zy9","https://ap.wps.com/l/cbCais2F0YAC7zy9","pdf",4149813,26,"English","# Abstract\n# Introduction\n# Method Overview\n## Prompt optimization and rejection sampling\n## Attack settings and evaluation\n# Findings and implications","[{\"question\":\"What impact does instruction-tuning have on memorization in LLMs, according to the paper?\",\"answer\":\"Instruction-tuning can increase memorization and the discoverability of pre-training data in aligned models, enabling instruction-based prompts to uncover higher memorization levels than direct prompting with training data.\"},{\"question\":\"How does the proposed black-box prompt optimization method work?\",\"answer\":\"An attacker LLM agent uses an iterative rejection-sampling loop to craft instruction-based prompts, minimizing prompt overlap with training data while maximizing overlap between the victim model’s outputs and training continuations via a reward function.\"},{\"question\":\"What are the two attack settings evaluated in the study?\",\"answer\":\"The paper evaluates an analytical setting to estimate an empirical upper bound, with and without access to responses for prompt initialization, and a practical classifier-based method to assess memorization without access to memorized data.\"}]","ALPACA AGAINST VICUNA - Using LLMs to Uncover Memorization of LLMs | PDF",9]