[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84789-en":3,"doc-seo-84789-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84789,5909877438554,"Maeve","https://ap-avatar.wpscdn.com/avatar/5600025385ad2bf12a7?_k=1778553567797529272",8,"Research & Report","Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies","Vision-language-action (VLA) models support robot navigation from natural language and visual goals but can be misled by perceptual distractions and ambiguous scene understanding. This paper evaluates visual grounding for VLA navigation policies and introduces a real-time segmentation-based approach that marks traversable regions in green and non-traversable regions in red using SegFormer. Experiments with OmniVLA on the Grand Tour dataset show 27–44% reduction in mean far-waypoint error, especially for long instructions, while image goals benefit minimally.","Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation  \nPolicies  \nAdrian Szvoren  \nDepartment of Computer Science University College London [adrian.szvoren.23@ucl.ac.uk](adrian.szvoren.23@ucl.ac.uk)  \nDimitrios Kanoulas  \nDepartment of Computer Science University College London [d.kanoulas@ucl.ac.uk](d.kanoulas@ucl.ac.uk)  \nNilufer Tuptuk  \nDepartment of Security and Crime Science University College London [n.tuptuk@ucl.ac.uk](n.tuptuk@ucl.ac.uk)  \narXiv :2607 .05 122v 1 [ cs .CV] 6 Jul 2026  \nAbstract—Vision-language-action (VLA) models enable robot navigation from natural language and visual goals, but remain susceptible to perceptual distractions and ambiguous scene interpretations. This paper presents the first empirical evaluation of visual grounding for VLA navigation policies. We propose areal-time segmentation-based grounding method that highlights traversable areas in green and non-traversable areas in red using SegFormer. Two variants are evaluated: observation-only segmentation and joint observation-goal augmentation. Using OmniVLA on the Grand Tour dataset, we show that visual grounding reduces the mean waypoint error by 27 − 44% atthe farthest waypoint, depending on the instruction length. The benefits are greater for long instructions than for short instructions, and grounding provides little improvement for image goals. Normalized error analysis indicates that grounding primarily acts as a trajectory length regularizer, reducing the predicted pathlength by 30% without improving per-unit-distance reasoning. Our results indicate that visual grounding offers a simple, computationally inexpensive method to improve VLA navigation without model retraining, although it cannot compensate for missing training signals in out-of-distribution instructions.  \nI. INTRODUCTION  \nVision-language-action (VLA) models have emerged as a promising development for robot navigation, enabling agents to interpret visual observations and natural language instructions to generate actions or navigation trajectories. Unlike traditional navigation systems that often rely on pre-built maps or explicit waypoint sequences, VLA models leverage largescale pretraining on diverse datasets to generalize to novel environments and instructions without task-specific fine-tuning. Recent advances have demonstrated impressive capabilities in following free-form language goals and adapting to unseen spatial layouts [9, 3, 4] . However, despite these successes, VLA policies remain susceptible to perceptual distractions, ambiguous scene interpretations, and the inherent difficulty of translating high-dimensional visual input into precise navigational actions.  \nOne promising direction for improving VLA robustness is visual grounding. Visual grounding techniques augmentor preprocess visual input to make spatial relationships and traversability cues more explicit to the underlying model. In the context of vision-language models, methods have shown  \nthat overlaying images with segmentation masks, numerical labels, or action visualizations can significantly enhance reasoning and visual question answering accuracy [19, 13] . This leads to the following research question: can visual grounding similarly benefit navigation VLA policies, where the agent must continuously reason about traversable space, obstacles, and goal-directed movement in real-time?  \nTo empirically evaluate this, we instantiate two variants of our grounding pipeline: one that segments only the current visual observation, and a second that additionally augments the goal modality (appending explicit traversability instructions to language goals or segmenting image goals) . We focus on three key constraints: real-time inference, generalization to unseen environments, and multimodal goal support.  \nThe key contributions of this paper are: (1) to our knowledge, the first empirical evaluation of visual grounding fora VLA navigation policy; (2) a real-time segmentati","cbCairfo4GnGxTWE","https://ap.wps.com/l/cbCairfo4GnGxTWE","pdf",2720610,1,7,"English","en",105,"# Introduction\n# Related Work\n## Vision-Language and Vision-Language-Action Models for Navigation\n# Method and Grounding Pipeline\n## Segmentation-based Visual Grounding and Goal Augmentation\n# Experimental Setup and Results\n## Evaluation Metrics and Waypoint Error Reduction\n# Limitations and Future Work","[{\"question\":\"What problem do the authors target in VLA navigation policies?\",\"answer\":\"They address susceptibility to perceptual distractions and ambiguous scene interpretations that can degrade accurate navigation from visual and language inputs.\"},{\"question\":\"How does the proposed visual grounding method work?\",\"answer\":\"It uses SegFormer to perform real-time semantic segmentation, overlaying traversable areas in green and non-traversable areas in red, then using these overlays to augment both observations and goals.\"},{\"question\":\"What improvements does grounding bring, and when is it most effective?\",\"answer\":\"On the Grand Tour dataset, grounding reduces mean waypoint error by 27–44% at the farthest waypoint, with larger gains for long instructions; it provides little improvement for image goals.\"}]",1784198246,18,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"green-for-go-red-for-no-visual-grounding-via-semantic-segmentation-for-vla-navigation-policies","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/green-for-go-red-for-no-visual-grounding-via-semantic-segmentation-for-vla-navigation-policies/84789/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem do the authors target in VLA navigation policies?","Question",{"text":75,"@type":76},"They address susceptibility to perceptual distractions and ambiguous scene interpretations that can degrade accurate navigation from visual and language inputs.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed visual grounding method work?",{"text":80,"@type":76},"It uses SegFormer to perform real-time semantic segmentation, overlaying traversable areas in green and non-traversable areas in red, then using these overlays to augment both observations and goals.",{"name":82,"@type":73,"acceptedAnswer":83},"What improvements does grounding bring, and when is it most effective?",{"text":84,"@type":76},"On the Grand Tour dataset, grounding reduces mean waypoint error by 27–44% at the farthest waypoint, with larger gains for long instructions; it provides little improvement for image goals.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]