[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85203-en":3,"doc-seo-85203-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85203,962075114101,"Seraphina","https://ap-avatar.wpscdn.com/avatar/e000253a75eb197efd?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780044092746381165",8,"Research & Report","Imperceptible and Reversible Adversarial Examples against Vision-Language Models for Privacy Protection","Vision–Language Models (VLMs) enable strong multi-modal understanding but also expose users to text-driven privacy attacks, where adversaries harvest online photos and query VLMs to infer sensitive attributes. Existing reversible adversarial example methods focus on vision-only settings and fail under multi-modal prompting, while current VLM attacks rely on high-frequency noise that harms visual fidelity. CloakDiff introduces reversible, high-fidelity privacy protection via diffusion-based adversarial editing and an invertible network enabling lossless recovery, with enhanced transferability through pixel embedding perturbations and latent cross-attention manipulation, further guided by EDM-Heuristic Sampling.","Imperceptible and Reversible Adversarial Examples against Vision-Language Models for Privacy Protection  \nQi Lu  \n[luqi@hust.edu.cn](luqi@hust.edu.cn)[ ](luqi@hust.edu.cn)School of Cyber Science and Engineering, Huazhong University of Science and Technology  \nZijing Li  \n[lizijing@hust.edu.cn](lizijing@hust.edu.cn)[ ](lizijing@hust.edu.cn)School of Software and engineering, Huazhong University of Science and Technology  \nZiqi Zhou  \n[zhouziqi@cqu.edu.cn](zhouziqi@cqu.edu.cn)[ ](zhouziqi@cqu.edu.cn)College of Computer Science, Chongqing University  \nYufei Song∗ [yufei17@hust.edu.cn](yufei17@hust.edu.cn)[ ](yufei17@hust.edu.cn)School of Cyber Science and Engineering, Huazhong University of Science and Technology  \nLulu Xue  \n[lluxue@hust.edu.cn](lluxue@hust.edu.cn)[ ](lluxue@hust.edu.cn)School of Cyber Science and Engineering, Huazhong University of Science and Technology  \nMinghui Li  \n[minghuili@hust.edu.cn](minghuili@hust.edu.cn)[ ](minghuili@hust.edu.cn)School of Software Engineering, Huazhong University of Science and Technology  \narXiv :2607 . 10329v1 [ cs .CV] 11 Jul 2026  \nShengshan Hu  \n[hushengshan@hust.edu.cn](hushengshan@hust.edu.cn)[ ](hushengshan@hust.edu.cn)School of Cyber Science and Engineering, Huazhong University of Science and Technology  \nLeo Yu Zhang  \n[leo.zhang@griffith.edu.au](leo.zhang@griffith.edu.au)[ ](leo.zhang@griffith.edu.au)School of Information and Communication Technology, Griffith University  \nAbstract  \nVision–Language Models (VLMs) offer powerful multi-modal ability but also expose users to text-based privacy attacks where adversaries crawl online photos and query VLMs to extract sensitive attributes. Existing reversible adversarial example (RAE) methods protect images in purely visual tasks but fail in multi-modal settings, and current adversarial examples on VLMs rely on high-frequency noise that severely degrades visual quality. We propose CloakDiff, the first framework for reversible, high-fidelity privacy protection against text-based query attacks in VLMs. CloakDiff produces imperceptible adversarial examples by combining diffusion-based adversarial editing with an invertible network that embeds the original image for lossless recovery. It perturbs both pixel-space embeddings and manipulates latent cross attention maps to ensure strong cross-model and cross-prompt transferability while preserving global visual structure. To further enhance fidelity, we design EDM-Heuristic Sampling, a principled diffusion schedule for adversarial guidance. Experiments on multiple datasets and VLMs demonstrate that CloakDiff delivers multi-modal privacy preservation with high visual quality and reversibility.  \nCCS Concepts  \n• Information systems → Multimedia information systems; • Security and privacy → Social aspects of security and privacy.  \nKeywords  \nAdversarial attack, Privacy protection, Diffusion Model  \n1 Introduction  \nVision-Language Models (VLMs) [17, 22, 49] exhibit strong visual feature extraction and understanding capabilities, and are widely  \n∗ Yufei Song is the corresponding author.  \nFigure 1: Illustration of privacy risks under VLMs.  \nused in downstream tasks such as caption generation [8] and visual question answering (VQA) [16]. However, when misused by malicious actors, they raise serious concerns about image privacy. Existing studies [39] show that attackers can crawl user images and exploit VLMs to launch large-scale automated text queries that extract sensitive information. Even benign-looking prompts can cumulatively expose private attributes such as gender, age, or location [38]. As shown in Fig. 1, an attacker exploits a photo the user shared for decoration advice, obtains cues like “a poster” and “aguitar,” and infers the user is in his twenties [45] .  \nTo defend against such attacks, a practical privacy protection method should preserve the image’s utility for legitimate users while preventing unauthorized models from inferring private cues. Since the threat arises fro","cbCairvScWlrT02S","https://ap.wps.com/l/cbCairvScWlrT02S","pdf",3193457,3,1,10,"English","en",105,"# Abstract\n# 1 Introduction\n## Privacy risks in VLMs\n## Reversible adversarial examples for privacy\n## Key challenge and proposed direction","[{\"question\":\"How does CloakDiff achieve reversible, high-fidelity privacy protection?\",\"answer\":\"CloakDiff uses diffusion-based adversarial editing combined with an invertible network to embed the original image for lossless recovery. It perturbs pixel-space embeddings and manipulates latent cross-attention maps to maintain global structure while improving cross-model and cross-prompt transferability, aided by an EDM-Heuristic Sampling diffusion schedule.\"}]",1784201732,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"imperceptible-and-reversible-adversarial-examples-against-vision-language-models-for-privacy-protection","",{"@graph":36,"@context":77},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/imperceptible-and-reversible-adversarial-examples-against-vision-language-models-for-privacy-protection/85203/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"How does CloakDiff achieve reversible, high-fidelity privacy protection?","Question",{"text":75,"@type":76},"CloakDiff uses diffusion-based adversarial editing combined with an invertible network to embed the original image for lossless recovery. It perturbs pixel-space embeddings and manipulates latent cross-attention maps to maintain global structure while improving cross-model and cross-prompt transferability, aided by an EDM-Heuristic Sampling diffusion schedule.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,126],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":22,"slug":125},"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":98,"slug":129},19,"General","general"]