[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82017-en":3,"doc-seo-82017-105":29,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":11,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82017,7971461740886,"Theodore","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks","Multimodal unlearning for multimodal foundation models addresses the need for selective removal of sensitive, copyrighted, biased, or unsafe cross-modal associations learned from web-scale training data. Retraining after deletion requests or policy updates is often impractical, and targeted forgetting is difficult because knowledge is distributed across shared representations. The survey systematizes vision, language, audio, and video unlearning approaches with a unified taxonomy, comparing deletion strength, retention, efficiency, reversibility, and robustness, and outlines open problems for deployment. A curated repository is released.","Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks  \nNobin Sarwar , Shubhashis Roy Dipta , Zheyuan Liu , Vaidehi Patil   University of Maryland, Baltimore County  University of Notre Dame  UNC Chapel Hill  \n{sms2, [sroydip1}@umbc.edu](sroydip1}@umbc.edu)  \n[zliu29@nd.edu](zliu29@nd.edu) , [vaidehi@cs.unc.edu](vaidehi@cs.unc.edu)  \narXiv :2607 .07907v 1 [ cs .LG] 8 Jul 2026  \nAbstract  \nWith the growing adoption of VLMs, DMs, LLMs, and AFMs, these multimodal foundation models can inadvertently encode sensitive, copyrighted, biased, or unsafe cross-modal associations that originate from their training data.  \nRetraining after deletion requests or policy updates is often impractical, and targeted forgetting remains difficult because knowledge is distributed across shared representations. Multimodal unlearning addresses this challenge by enabling selective removal across modalities while retaining overall utility. This survey offers a unified, system-oriented view of multimodal unlearning across vision, language, audio, and video, grounded in recent advances, emerging applications, and open problems. Our taxonomy enables systematic comparison across model architectures and modalities, clarifying trade-offs among deletion strength, retention, efficiency, reversibility, and robustness. This survey highlights open problems and practical considerations to support future research and deployment of multimodal unlearning. We release a curated repository.1  \n1 Introduction  \nMultimodal foundation models, including Vision Language Models (VLMs), Diffusion Models (DMs), Large Language Models (LLMs) and Audio Foundation Models (AFMs)-based (Ho et al., 2020 ; Team et al., 2023 ; Yang et al., 2025a ; Chu et al., 2023 ; Huang et al., 2024c) generators, support image, text, video, and audio understanding and generation at scale. Training on web-scale multimodal data improves generalization, but it can also induce memorization and undesired associations involving sensitive, copyrighted, biased, or unsafe content across modalities. As a result, deployed models may need to forget specific items  \n1 [https://smsnobin77.github.io/](https://smsnobin77.github.io/)[ ](https://smsnobin77.github.io/)Awesome-Multimodal-Unlearning/  \nUnlearning Intervention Points  \nFigure 1: Unlearning intervention points for a Multimodal Foundation Model (MFM) . Methods intervene at the data side, during training, via architectureconstrained edits, or at decoding time, producing an updated model (MFM′) with reduced influence from targeted content. Training-free methods use closed-form parameter or representation edits (denoted by ∆) to directly transform the model without retraining.  \n\n| Survey | Venue & Year | System-first | Text | Image | Video | Audio |\n| --- | --- | --- | --- | --- | --- | --- |\n| Si et al., 2023 | arXiv’23 |  | ✔ |  |  |  |\n| Liu et al., 2024f | arXiv’24 | ✔ | ✔ | ✔ |  |  |\n| Blanco-Justicia et al., 2025 | AIR’25 | ✔ | ✔ |  |  |  |\n| Liu et al., 2025b | NMI’25 |  | ✔ |  |  |  |\n| Feng et al., 2025b | arXiv’25 | ✔ | ✔ | ✔ |  | ✔ |\n| Geng et al., 2025 | arXiv’25 | ✔ | ✔ | ✔ |  |  |\n| Ours | ACL’26 | ✔ | ✔ | ✔ | ✔ | ✔ |\n\nTable 1: Comparison of multimodal unlearning surveys across modalities and system-first taxonomy coverage.  \nor concepts, such as a copyrighted artwork, a private face, or a harmful trope, while retaining performance on the remaining data (Fan et al., 2023 ; Gandikota et al., 2023 ; Zhang et al., 2024d ; Sun et al., 2024 ; Chen et al., 2025d,b ; Facchiano et al., 2025) . When deletion requests or policy updates affect only part of the training signal, retraining from scratch is often impractical (Voigt and Von dem Bussche, 2017 ; Goldman, 2020) . Targeted removal is challenging because knowledge is distributed in shared representations, so eliminating one association can disrupt unrelated behavior.  \nThese challenges have driven growing interest in multimodal unlearning as a mec","cbCairL6NhxOi5x6","https://ap.wps.com/l/cbCairL6NhxOi5x6","pdf",2768855,1,29,"English","en",105,"# Introduction\n# Unlearning Intervention Points\n# Data-Side Interventions\n## Data-Path Perturbation Unlearning\n## Data Hygiene and Prompt Normalization\n# Training-Time Edits\n## Direct Gradient\n## Constrained Updates\n## Mask-Driven Selective Unlearning\n## Distillation-Based\n# Architecture Constrained\n# Training-Free Unlearning\n## Weight-space Linear Unlearning\n## Representation Projection\n# Decoding Time\n## Guidance-Path Control\n## Conditioning-Path Control","[{\"question\":\"How does the survey organize multimodal unlearning methods across modalities?\",\"answer\":\"It provides a unified, system-oriented taxonomy spanning vision, language, audio, and video, comparing approaches by deletion strength, retention, efficiency, reversibility, and robustness.\"}]",1784177603,73,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":27},"multimodal-unlearning-across-vision-language-video-and-audio-survey-of-methods-datasets-and-benchmarks","",{"@graph":35,"@context":77},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/multimodal-unlearning-across-vision-language-video-and-audio-survey-of-methods-datasets-and-benchmarks/82017/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-08-01","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":11},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"How does the survey organize multimodal unlearning methods across modalities?","Question",{"text":75,"@type":76},"It provides a unified, system-oriented taxonomy spanning vision, language, audio, and video, comparing approaches by deletion strength, retention, efficiency, reversibility, and robustness.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":45,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":45,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":45,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":45,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":98,"slug":130},19,"General","general"]