[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86528-en":3,"doc-seo-86528-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86528,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Think When It Matters Conditional VLM Reasoning for Social Navigation with RL Policies","As mobile robots move into everyday human environments, social navigation must guarantee comfort, safety, and trust. Existing reinforcement learning (RL) policies support fast reactive control but lack flexible semantic reasoning and often generalize poorly to complex social situations. Vision-language models (VLMs) improve context understanding but introduce high compute cost and slow inference, hindering real-time use. HUMA combines an efficient reactive RL policy with selectively conditioned VLM reasoning when humans enter sensitive proximity zones, improving success on Social-MP3D and Social-HM3D while reducing personal space violations and collisions.","arXiv :2607 . 10991v1 [ cs .RO] 13 Jul 2026  \nThink When It Matters: Conditional VLM Reasoning for Social Navigation with RL Policies  \nAli Ahmadi, Hamed Rahimi, Adrien Jacquet Crtides,  \nMarie Samson, Mahdi Khoramshahi, Mohamed Chetouani  \nInstitut des Systmes Intelligents et de Robotique (ISIR) Sorbonne Universit France  \n{[lastname](lastname}@isir.upmc.fr)[}](lastname}@isir.upmc.fr)[@isir.upmc.fr](lastname}@isir.upmc.fr)  \nAbstract:  \nAs mobile robots become more integrated into everyday human environments, social robot navigation is becoming essential for ensuring human comfort, safety, and trust. While reinforcement learning (RL) navigation policies provide the fast inference and reactive behavior necessary for real-time deployment, they still lack flexible semantic reasoning capabilities and often fail to generalize to complex social scenarios. Recent approaches have increasingly turned to vision-language models (VLMs) in place of RL policies to improve semantic and social reasoning in robot navigation. Nevertheless, their high computational cost and slow inference remain major barriers to real-time deployment. To overcome these limitations, we introduce HUMA (Hybrid Understanding for Multi-modal social Navigation), a hybrid architecture that dynamically balances the computational efficiency of RL policies with the deep semantic understanding of VLMs. Our approach uses a reactive RL policy to handle low-density, routine navigation tasks, while conditioning it on a post-trained high-level VLM when a human enters sensitive situations, such as the robot’s proximity zone. We evaluate HUMA on the Social-MP3D and Social-HM3D benchmarks, where it achieves task success improvements of 20% and 3%, respectively, while significantly reducing personal space violations and human collisions against state-of-the-art baselines. Extensive ablation studies validate each architectural component, and real-world deployment on the Miroka¨ı mobile robot further demonstrates the practical viability of our approach.  \n1 Introduction  \nAs mobile robots increasingly transition from isolated settings into dynamic, human-centric environments such as airports or hospitals, social robot navigation appears as one of the main challenges of Human-Robot Interactions (HRI) in modern societies [1, 2] . Social Navigation (also referred to as Human-aware or Socially-aware Navigation) lies at the intersection of Human-Robot Interaction (HRI) and Robot Motion Planning [3, 1] . It concerns the ability of robots to navigate safely and efficiently in human-populated environments while respecting social norm and human safety such as judging a path’s clearance relative to human movements and obstacles [4] . This requires an agent to dynamically read social cues and anticipate human intent, transforming classical Motion Planning into a socially-aware optimization task. By adapting its trajectory to maintain a non-disruptive presence, a socially aware robot allows for safe, intuitive and comfortable interactions in shared spaces [5] . While early frameworks relied on fixed geometric rules, current approaches to social navigation have leveraged learning-based paradigms to tackle complex environments [6, 7] . Specifically, policies trained with reinforcement learning (RL) have become a standard approach, largely due to their fast inference and reactive performance in real-time deployment. However, although RL methods are often computationally intensive and data-hungry during training, modeling the nuances of human social behavior typically requires large-scale datasets, and purely RL-based approaches frequently  \nFigure 1: Overview. Existing approaches for social navigation face a fundamental trade-off: RLbased policies (left) offer fast, reactive inference but lack semantic reasoning for complex social scenarios, while VLM-based methods (middle) provide rich contextual understanding at the cost of high computational overhead and slow inference, preventing real-time dep","cbCaisD3nx8JHGkx","https://ap.wps.com/l/cbCaisD3nx8JHGkx","pdf",6055438,3,1,15,"English","en",105,"# 1 Introduction\n## Problem: RL vs VLM trade-off\n## Core idea: selective VLM arbitration\n## Paper contribution: HUMA overview","[{\"question\":\"Why do reinforcement learning navigation policies struggle in social scenarios?\",\"answer\":\"They provide fast reactive control but lack flexible semantic reasoning and may generalize poorly to rare or complex human interactions.\"},{\"question\":\"What problem do vision-language models face for real-time robot navigation?\",\"answer\":\"Their computational overhead and slow inference limit deployment for real-time social navigation.\"},{\"question\":\"How does HUMA decide when to use VLM reasoning instead of RL-only control?\",\"answer\":\"HUMA uses a reactive RL policy for routine low-density navigation and activates a post-trained high-level VLM when a human enters the robot’s proximity zone, especially in sensitive situations.\"}]",1784212416,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"think-when-it-matters-conditional-vlm-reasoning-for-social-navigation-with-rl-policies","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/think-when-it-matters-conditional-vlm-reasoning-for-social-navigation-with-rl-policies/86528/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do reinforcement learning navigation policies struggle in social scenarios?","Question",{"text":75,"@type":76},"They provide fast reactive control but lack flexible semantic reasoning and may generalize poorly to rare or complex human interactions.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What problem do vision-language models face for real-time robot navigation?",{"text":80,"@type":76},"Their computational overhead and slow inference limit deployment for real-time social navigation.",{"name":82,"@type":73,"acceptedAnswer":83},"How does HUMA decide when to use VLM reasoning instead of RL-only control?",{"text":84,"@type":76},"HUMA uses a reactive RL policy for routine low-density navigation and activates a post-trained high-level VLM when a human enters the robot’s proximity zone, especially in sensitive situations.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]