[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85338-en":3,"doc-seo-85338-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85338,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","Technical Report on the CVPR 2026@AdvML Workshop Challenge","Vision-language agents (VLAs) are increasingly applied to interpret complex driving scenes and support safety-critical reasoning. This report presents the CVPR 2026@AdvML Workshop Challenge on adversarial multimodal attacks against autonomous-driving VLAs. Scenes are represented by six synchronized camera images and structured driving QA pairs. Participants craft adversarial images and suffix-only textual perturbations to deviate from reference answers while preserving fidelity and limiting text cost. The two-phase competition includes a hidden black-box model in Phase II to evaluate transferability, followed by task design, rules, evaluation, and analysis of five leading submissions.","arXiv :2607 . 11560v1 [ cs .CV] 13 Jul 2026  \nTechnical Report on the CVPR 2026@AdvML Workshop Challenge  \nTianyuan Zhang 1 ,†, Zonglei Jing 1 ,†, Jiangfan Liu 1 ,†, Ligong Zhang2 ,‡, Ke Ma2 ,‡, Chengzhi Sun2 ,‡ Xiaohai Xu2 ,‡, Zhirui Zhang2 ,‡, Qianqian Xu3 ,‡, Qingming Huang2 ,‡, Hanyu Fang4 ,‡, Junhua Liu5 ,6 ,‡ Zheng Wang4 ,‡, Xiaoliang Liu7 ,‡, Yuanbo Li8 ,‡, Shuai Gui8 ,‡, Bin Wang8 ,‡, Menghe Zheng8 ,‡, Jing Nie8 ,‡ Hanyang Meng8 ,‡, Zeyang Zhang8 ,‡, Xiang Zhang8 ,‡, Yongxuan Zhu9 ,‡, Rui Ding 10 ,‡, Hainan Li 11 ,† Yongkang Zhang 1 ,†, Zhilei Zhu 11 ,†, Xianglong Kong 11 ,†, Jin Hu 1 , 12 ,†, Zonghao Ying 1 ,†, Yisong Xiao 1 ,† Lei Chen 13 ,† Haotong Qin 14 ,†, Jiakai Wang 12 ,†, Aishan Liu 1 ,†, Ruikai Li 1 ,†, Julia Karbing 15 ,†, Yinpeng Dong 13 ,† Zhenfei Yin 15 ,†, Shao Jing 16 ,†, Xia Hu 16 ,†, Jingyi Xu 1 ,†, Juntao Dai 17 ,†, Xinyun Chen 18 ,†, Vishal M. Patel 19 ,† Xianglong Liu 1 , 12 ,†, Dawn Song20 ,†, Alan Yuille 19 ,†, Philip H. S. Torr15 ,†, Dacheng Tao21 ,†  \n1 Beihang University 2 University of Chinese Academy of Sciences 3 Institute of Computing Technology, Chinese Academy of Sciences 4 Tongji University  \n5 iFLYTEK Co., Ltd. 6 Anhui Laboratory for Safe Artificial Intelligence in the Yangtze River Delta 7 Wenzhou Business College 8 Jiangnan University  \n9 Guangzhou City University of Technology 10 Inceptio Technology 11 Institute of Dataspace 12 Zhongguancun Laboratory 13 Tsinghua University  \n14 ETH Z¨urich 15 University of Oxford 16 Shanghai AI Laboratory 17 BAAI 18 Meta 19 Johns Hopkins University  \n20 University of California, Berkeley 21 Nanyang Technological University  \n† Organizer ‡ Challenger  \n[https://cvpr26-advml.github.io/](https://cvpr26-advml.github.io/)  \nAbstract  \nVision-language agents (VLAs) are increasingly used to interpret complex driving scenes and support safety-critical reasoning. This report presents the CVPR 2026@AdvML Workshop Challenge on adversarial multimodal attacks against autonomous-driving VLAs. Built on DriveLM-style multi-view visual question answering, the challenge represents each scene with six synchronized camera images anda structured collection of driving-related question-answer pairs. Participants generate adversarial images and suffixonly textual perturbations that induce model responses to deviate from reference answers while preserving image fidelity and limiting textual cost. The competition comprises two phases, with Phase II adding a hidden black-box model to assess transferability. We describe the task design, submission rules, evaluation protocol, and leaderboard results, and then examine five leading submissions for which technical reports were available. Across these reports, several recurring patterns emerge: image-side attacks are favored by the suffix penalty; scene-level, multi-view optimization is more effective than treating views in isolation; QA types and graph structure provide useful priors for allocating attack budget; feature-space objectives can improve black-box transfer; and typographic content embedded in camera images exposes a persistent vulnerability in driv-  \ning VLAs. These findings provide a practical reference for future robustness evaluation and defense design in multimodal autonomous-driving systems.  \n1. Introduction  \nRecent multimodal foundation models and vision-language agents have demonstrated strong capabilities in perceiving traffic scenes, reasoning about objects and their relations, and answering decision-oriented driving questions [1, 7, 11, 14, 20] . In autonomous driving, however, these capabilities also create a safety-critical attack surface. A model that misinterprets a traffic light, overlooks a nearby pedestrian, or produces an unsafe response to a planning question may compromise downstream decision making. The CVPR 2026 @ AdvML Workshop Challenge investigates this risk through adversarial multimodal attacks against driving VLAs, building on the broader literature on adversarial examples and robustness","cbCaivewou5hFsgT","https://ap.wps.com/l/cbCaivewou5hFsgT","pdf",2466993,3,1,21,"English","en",105,"# 1. Introduction\n# 2. Competition Overview\n## 2.1. Theme and Task","[{\"question\":\"What is the CVPR 2026@AdvML Workshop Challenge about?\",\"answer\":\"It evaluates adversarial multimodal attacks against autonomous-driving vision-language agents in safety-critical scenarios using a DriveLM-style visual question answering setup.\"},{\"question\":\"How are driving scenes and inputs structured for the challenge?\",\"answer\":\"Each scene is represented by six synchronized camera views plus a structured collection of driving-related question-answer pairs.\"},{\"question\":\"What is added in Phase II of the competition?\",\"answer\":\"Phase II introduces a hidden black-box model so transferability becomes a central objective in addition to attack success.\"}]",1784202607,53,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"technical-report-on-the-cvpr-2026advml-workshop-challenge","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/technical-report-on-the-cvpr-2026advml-workshop-challenge/85338/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the CVPR 2026@AdvML Workshop Challenge about?","Question",{"text":75,"@type":76},"It evaluates adversarial multimodal attacks against autonomous-driving vision-language agents in safety-critical scenarios using a DriveLM-style visual question answering setup.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How are driving scenes and inputs structured for the challenge?",{"text":80,"@type":76},"Each scene is represented by six synchronized camera views plus a structured collection of driving-related question-answer pairs.",{"name":82,"@type":73,"acceptedAnswer":83},"What is added in Phase II of the competition?",{"text":84,"@type":76},"Phase II introduces a hidden black-box model so transferability becomes a central objective in addition to attack success.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]