[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84807-en":3,"doc-seo-84807-105":29,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84807,2336464648322,"Aria","https://ap-avatar.wpscdn.com/avatar/2200025388227c56fec?_k=1778556882303663488",8,"Research & Report","Evaluating and Understanding Model Editing for Medical Vision Language Models","Model editing offers a fast, targeted approach to fix postdeployment errors in medical vision–language models (VLMs) without costly retraining. Existing multimodal model-editing benchmarks emphasize general-purpose settings that fail to capture clinical domain variability and requirements. M3 Bench is introduced as a clinically grounded benchmark that tests reliability, precision, locality, generalizability, modality/text variation, protocol shifts, knowledge composition, and temporal progression via 16,276 questions and both single and sequential edits.","arXiv :2607 .053 10v 1 [ cs .AI] 6 Jul 2026  \nEvaluating and Understanding Model Editing for Medical Vision Language Models  \nGuli Zhu* 1 , Chenwei Wu* 1 , and Liyue Shen 1  \nEECS, University of Michigan, Ann Arbor, MI 48105, USA {gulizhu, chenweiw, [liyues}@umich.edu](liyues}@umich.edu)  \nAbstract. Model editing promises a fast, targeted way to correct postdeployment mistakes in medical vision–language models (VLMs) without costly retraining. However, existing multimodal model editing benchmarks focus on general-purpose tasks and do not reflect realistic clinical domain requirements and variability. To address this, we introduce M3 Bench, a clinically grounded benchmark for multimodal model editing that evaluates whether an edit remains reliable, precise and generalizable under the challenges of image and text variation, modality and protocol shifts, clinical knowledge composition, and temporal progression. M3 Bench contains 16,276 questions spanning diverse anatomy, modalities, and specialties, and supports both single and sequential edits. By evaluating 4 representative editors across 6 medical and general VLMs, we indicate that no method excels across all criteria. Gradientbased editors achieve strong transfer but suffer from catastrophic locality violations, whereas memory-based methods preserve locality but lack compositional generality and exhibit high backbone-dependent hyperparameter sensitivity. We further attribute these failures to the latent space geometry of VLMs and how different editing methods shift its landscape. Overall, M3 Bench establishes a rigorous clinical stress test for multimodal model editing and offers actionable guidance for safer post-deployment adaptation. The benchmark is publicly available at [https://github.com/BioMed-AI-Lab-U-Michgan/M3Bench](https://github.com/BioMed-AI-Lab-U-Michgan/M3Bench).  \nKeywords: Model editing · Medical VLMs · Benchmark  \n1 Introduction  \nVision–Language Models (VLMs) hold great promise in supporting multimodal clinical workflows such as automated radiology report generation and real-time surgical assistance [1, 8] . However, as shown in Fig. 1a, real-world deployment of these large-scale models is not a one-time milestone: once a model enters clinical practice, it is exposed to continuously changing patient data and clinical protocols that may not be observed during model development. As a result,  \neven well-trained models will inevitably make errors after deployment, potentially leading to serious consequences, such as failing to recognize rare diseases,* Equal contribution, order determined by coin flip.  \n2 G. Zhu and C. Wu et al.  \n(a) (b)  \nFig. 1: Post-deployment workflow and model editing performance. (a) Postdeployment workflow of multimodal model editing with clinician-in-the-loop feedback. Left: model development stage with extensive model training. Right: model deployment stage with lightweight and targeted model editing to correct errors on the fly. (b) Editing performance radar (sequential-edit) summarizing performance across our M3 Bench evaluation dimensions for LLaVA-Med-v1 .5-Mistral-7b.  \nmisinterpreting complex patient cases involving compositional conditions, or incorrectly analyzing studies acquired with new scanners [2, 15, 22, 32] .  \nAddressing such post-deployment errors through global model updates, such as full fine-tuning or retraining, is often impractical, as they are computationally expensive, data-intensive, time-consuming, and may degrade prior model capabilities [26] . Model Editing has therefore emerged as a promising alternative for targeted intervention [28], enabling precise corrections by updating a small subset of model parameters or incorporating external parameter memory, while keeping the rest of the model unchanged. While these techniques have been increasingly studied for general-domain LLMs and VLMs [9], how model editing behaves under realistic multimodal clinical tasks remains largely underexplored. Editing methods ","cbCain9Au6JMRrt1","https://ap.wps.com/l/cbCain9Au6JMRrt1","pdf",8326188,1,33,"English","en",105,"# Introduction\n## Post-deployment errors and limitations of global updates\n## Reliability, Locality, and Generality in existing benchmarks\n## Clinical reformulation via M3 Bench\n# Benchmark design and evaluation settings\n## Clinically motivated variation types\n## Temporal consistency and longitudinal comparisons\n# Experiments and results\n## Editors across medical and general VLM backbones\n## Failure analysis via latent space geometry","[{\"question\":\"What do the authors find about different model editing methods across medical and general VLMs?\",\"answer\":\"No method satisfies all criteria: gradient-based editors transfer well but can cause catastrophic locality violations, while memory-based methods better preserve locality but may lack compositional generality and show high sensitivity to backbone-dependent hyperparameters.\"}]",1784198373,83,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":27},"evaluating-and-understanding-model-editing-for-medical-vision-language-models","",{"@graph":35,"@context":77},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/evaluating-and-understanding-model-editing-for-medical-vision-language-models/84807/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"What do the authors find about different model editing methods across medical and general VLMs?","Question",{"text":75,"@type":76},"No method satisfies all criteria: gradient-based editors transfer well but can cause catastrophic locality violations, while memory-based methods better preserve locality but may lack compositional generality and show high sensitivity to backbone-dependent hyperparameters.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":45,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":45,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":45,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":45,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":98,"slug":130},19,"General","general"]