[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82356-en":3,"doc-seo-82356-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},82356,687197207919,"Theodora","https://ap-avatar.wpscdn.com/avatar/a000253d6f5f7c60be?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779446848396160552",8,"Research & Report","Robustifying Vision-Language Models via Test-Time Prompt Adaptation","Pre-trained Vision-Language Models (VLMs) such as CLIP perform well in zero-shot settings, yet suffer severe degradation under adversarial perturbations. Existing test-time adaptation methods often depend on sample-level confidence heuristics, ignoring the underlying distributional structure and thus mixing confident adversarial errors with genuine semantic consistency. The work shows adversarial distortions are structurally brittle: augmented views preserve semantic integrity in their feature distributions. It proposes RITA, using optimal transport to align augmented visual feature distributions with textual prototypes, plus a dynamic cache for online refinement. Experiments demonstrate improved adversarial robustness without harming clean accuracy.","Robustifying Vision-Language Models via Test-Time Prompt Adaptation  \nXingyu Zhu 1 2 Huanshen Wu 2 Shuo Wang 2 * Beier Zhu 2 Jiannan Ge 2 Jiaheng Zhang 1 Long Chen 3  \narXiv :2607 .09450v1 [ cs .CV] 10 Jul 2026  \nAbstract  \nPre-trained Vision-Language Models (VLMs) such as CLIP achieve strong zero-shot generalization, but their performance degrades sharply under adversarial perturbations. Existing test-time adaptation methods typically rely on sample-level confidence heuristics, overlooking the intrinsic distributional structure of the data. This samplecentric approach limits robustness, as it fails to distinguish confident adversarial mispredictions from true semantic consistency. In this work, we observe that adversarial distortion is structurally brittle: while holistic representations are corrupted, semantic integrity is often preserved in the distribution of augmented views. Motivated by this insight, we propose RITA, a Robust testtIme prompT Adaptation framework that shifts from sample-level estimates to distribution-level alignment. Specifically, RITA employs optimal transport to align the distribution of augmented visual features with textual prototypes, mitigating adversarial outliers and rectifying cross-modal semantic misalignment. Furthermore, we introduce a dynamic cache to progressively accumulate reliable cues from the test stream for online refinement. Extensive experiments demonstrate that RITA significantly improves adversarial robustness without compromising clean accuracy.  \n1. Introduction  \nVision-Language Models (VLMs) (Li et al., 2022 ; Alayracet al., 2022 ; Li et al., 2023 ; Zhu et al., 2024b ; 2026c) like CLIP (Radford et al., 2021), pre-trained on massive imagetext pairs, have achieved remarkable zero-shot generalization. Despite this success, VLMs remain highly vulnerable to adversarial perturbations: imperceptible noise can  \n1National University of Singapore 2University of Science and Technology of China 3The Hong Kong University of  \n*  \nScience and Technology. Correspondence to: Shuo Wang \u003C[shuowang.edu@gmail.com](shuowang.edu@gmail.com) >.  \nProceedings of the 43 rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026 . Copyright 2026 by the author(s) .  \nFigure 1. Augmented views retain more semantic cues under adversarial perturbations, enabeling a cache for distribution alignment that improves adversarial performance. (a) Visualization of adversarially perturbed images, where each point represents an image and different colors denote ground-truth classes. (b) Visualization of multiple augmented views generated from the same adversarial image, colored by class label, where semantic structure partially re-emerges with improved class separability compared to (a) . (c) Our method leverages the selected augmented views as a cache and aligns them with textual prompts. (d) Performance comparison across different VLM backbones, demonstrating improved robustness under adversarial attacks.  \ncause severe performance degradation (Szegedy et al., 2014 ; Madry et al., 2018 ; Zhu et al., 2024a ; 2025b), posing security risks in real word applications.  \nExisting efforts to enhance the robustness of VLMs generally fall into two categories. The first line of work utilizes adversarial training (Mao et al., 2023 ; Schlarmann et al., 2024 ; Wang et al., 2024 ; Zhang et al., 2024), which aims to immunize models by explicitly integrating adversarial examples into the optimization loop. While effective, these approaches typically incur prohibitive computational costs due to on-the-fly attack generation and require access to taskspecific labeled data, thereby undermining the scalability and zero-shot flexibility inherent to foundation models. The second line of work explores test-time prompt tuning (Yoon et al., 2024 ; Shu et al., 2022 ; Zhao et al., 2025), an efficient paradigm that adapts learnable prompt contexts or predictions during inference without modifying model parameters. How","cbCaiiEwASKblCIo","https://ap.wps.com/l/cbCaiiEwASKblCIo","pdf",2602016,4,1,16,"English","en",105,"# Introduction\n## Robustness limits of vision-language models\n## Existing adversarial training and test-time prompt tuning\n## Key observation: augmented-view distributional structure\n## RITA approach: distribution-level alignment","[{\"question\":\"Why do current test-time adaptation methods struggle under adversarial attacks?\",\"answer\":\"They largely rely on sample-level confidence heuristics (e.g., entropy) that treat augmentations as independent points. This overlooks distributional structure, making it hard to separate confident adversarial mispredictions from true semantic consistency.\"},{\"question\":\"What is the core observation behind RITA?\",\"answer\":\"Adversarial distortions are structurally brittle across test-time augmentations. While holistic representations are corrupted, semantic integrity often remains preserved within the distribution formed by augmented views.\"},{\"question\":\"How does RITA improve robustness during inference?\",\"answer\":\"RITA aligns the distribution of augmented visual features with textual prototypes using optimal transport, mitigating adversarial outliers and correcting cross-modal semantic misalignment. It also uses a dynamic cache to accumulate reliable cues for online refinement.\"}]","Robustifying Vision-Language Models via Test-Time Prompt Adaptation | PDF",1784179847,40,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"robustifying-vision-language-models-via-test-time-prompt-adaptation","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":20},"https://docshare.wps.com/document/robustifying-vision-language-models-via-test-time-prompt-adaptation/82356/",{"url":53,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-29","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why do current test-time adaptation methods struggle under adversarial attacks?","Question",{"text":76,"@type":77},"They largely rely on sample-level confidence heuristics (e.g., entropy) that treat augmentations as independent points. This overlooks distributional structure, making it hard to separate confident adversarial mispredictions from true semantic consistency.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What is the core observation behind RITA?",{"text":81,"@type":77},"Adversarial distortions are structurally brittle across test-time augmentations. While holistic representations are corrupted, semantic integrity often remains preserved within the distribution formed by augmented views.",{"name":83,"@type":74,"acceptedAnswer":84},"How does RITA improve robustness during inference?",{"text":85,"@type":77},"RITA aligns the distribution of augmented visual features with textual prototypes using optimal transport, mitigating adversarial outliers and correcting cross-modal semantic misalignment. It also uses a dynamic cache to accumulate reliable cues for online refinement.","https://schema.org",{"og:url":53,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":30,"slug":119},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":47,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":47,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":47,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":47,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]