[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86122-en":3,"doc-seo-86122-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86122,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","AdvNav Behavior-Guided Black-Box Adversarial Attacks on Vision-Language Navigation","Vision-and-Language Navigation (VLN) systems remain vulnerable to adversarial visual disturbances, especially when deployed as practical black-box agents. Existing white-box approaches depend on gradient access and recursive backpropagation, making them computationally expensive and often unrealistic. Prior black-box attacks mainly address single-step decisions and fail to capture VLN’s multi-step sequential perception–action dependencies. AdvNav introduces a gradient-free, behavior-guided framework that perturbs first-person views using only observable inputs and outputs, with dual-granularity feedback for trajectory and action risk to iteratively find disruptive noise configurations. Experiments on R2R show strong attack success against HAMT and MapGPT variants.","AdvNav: Behavior-Guided Black-Box Adversarial Attacks on  \nVision-Language Navigation  \nChenyang Li  \nThe Hong Kong University of Science and Technology (Guangzhou) Guang Zhou, China [lichenyang20020820@gmail.com](lichenyang20020820@gmail.com)  \nKaige Li  \nSun Yat-sen University ShenZhen, China [likg@mail.sysu.edu.cn](likg@mail.sysu.edu.cn)  \nZeyu Jiang  \nThe Hong Kong University of Science and Technology (Guangzhou) Guang Zhou, China [zjiang739@gmail.com](zjiang739@gmail.com)  \nChanghao Chen∗ The Hong Kong University of Science and Technology (Guangzhou) GuangZhou, China [changhaochen@hkust-gz.edu.cn](changhaochen@hkust-gz.edu.cn)  \narXiv :2607 . 1 1063v 1 [ cs .AI] 13 Jul 2026  \nAbstract  \nDespite progress in Embodied AI, Vision-and-Language Navigation (VLN) systems remain vulnerable to adversarial visual disturbances. Most existing methods rely on white-box access to target model gradients, which is often unrealistic for real-world deployed systems and computationally exhaustive due to recursive backpropagation for optimization, limiting their applicability. While previous black-box methods predominantly target single-step, instantaneous decision tasks, they struggle to handle the task complexities and temporal dependencies of VLN. This highlights the need for a gradient-free attack method that can effectively disrupt the multi-step sequential perception–action loop using only observable inputs and outputs. Therefore, we propose AdvNav, a behavior-guided black-box adversarial attack framework that disturbs an agent’s first-person views during navigation. To construct an informative surrogate objective for effective optimization guidance in gradient-free search under the black-box setting, we design a dual-granularity behavior-based feedback, aggregating a trajectory-level performance score representing overall navigation degradation, an action-level reward score considering the potential decision risk, and a deviation indicator, all of which are extracted from the agent’s self-output behaviors. This feedback guides a hybrid optimization strategy that (i) heuristically tunes perturbation strength via adaptive updates and (ii) evolves noise spatial structure genetically, to iteratively discover the most disruptive noise configuration. Evaluated against Transformer-based model HAMT and LLM-based model MapGPT with two types of backbones on R2Rdataset, AdvNav achieves 49 . 70%, 65 . 96%, and 87 .30% Attack Success Rate, respectively. The result demonstrates the effectiveness  \n∗ Corresponding author.  \nPermission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission [and/or a fee. Request permissions from permissions@acm.org](and/or a fee. Request permissions from permissions@acm.org).  \nMM’26, Rio de Janeiro, Brazil  \n© 2026 Copyright held by the owner/author(s) . Publication rights licensed to ACM. ACM ISBN 978-1-4503-XXXX-X/2018/06  \n[https://doi.org/XXXXXXX.XXXXXXX](https://doi.org/XXXXXXX.XXXXXXX)  \nand generality of AdvNav, reveals critical perception vulnerabilities and offers insights for the design of future resilient VLN models.  \nCCS Concepts  \n• Computing methodologies → Vision for robotics; • Security and privacy; • Information systems → Multimedia information systems;  \nKeywords  \nVision-and-Language Navigation, Adversarial Attack, Black-box  \nACM Reference Format:  \nChenyang Li, Kaige Li, Zeyu Jiang, and Changhao Chen. 2026. AdvNav: Behavior-Guided Black-Box Adversarial Attacks on Vision-Language Navigation. In Proceedings of the 34th ACM International Conferenc","cbCaivM824SpsYyV","https://ap.wps.com/l/cbCaivM824SpsYyV","pdf",2259408,5,1,10,"English","en",105,"# Introduction\n## Vulnerability of VLN and need for black-box attacks\n# AdvNav Approach\n## Behavior-guided dual-granularity feedback\n## Hybrid gradient-free optimization strategy\n# Experiments\n## Results on R2R with HAMT and MapGPT","[{\"question\":\"Why do existing white-box adversarial methods have limited applicability to VLN?\",\"answer\":\"They require gradient access and rely on recursive backpropagation for optimization, which is computationally intensive and often unavailable for deployed proprietary systems.\"},{\"question\":\"How does AdvNav operate in a black-box setting?\",\"answer\":\"AdvNav is gradient-free and uses observable inputs and outputs, perturbing an agent’s first-person views while guiding search through behavior-derived feedback signals.\"},{\"question\":\"What feedback signals does AdvNav use to optimize adversarial noise?\",\"answer\":\"It combines trajectory-level performance scores, action-level reward scores reflecting potential decision risk, and a deviation indicator, extracted from the agent’s self-output behaviors.\"}]",1784208653,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"advnav-behavior-guided-black-box-adversarial-attacks-on-vision-language-navigation","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/advnav-behavior-guided-black-box-adversarial-attacks-on-vision-language-navigation/86122/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why do existing white-box adversarial methods have limited applicability to VLN?","Question",{"text":76,"@type":77},"They require gradient access and rely on recursive backpropagation for optimization, which is computationally intensive and often unavailable for deployed proprietary systems.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does AdvNav operate in a black-box setting?",{"text":81,"@type":77},"AdvNav is gradient-free and uses observable inputs and outputs, perturbing an agent’s first-person views while guiding search through behavior-derived feedback signals.",{"name":83,"@type":74,"acceptedAnswer":84},"What feedback signals does AdvNav use to optimize adversarial noise?",{"text":85,"@type":77},"It combines trajectory-level performance scores, action-level reward scores reflecting potential decision risk, and a deviation indicator, extracted from the agent’s self-output behaviors.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":22,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":20,"slug":137},19,"General","general"]