[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84938-en":3,"doc-seo-84938-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84938,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","UniLM-Nav: A Unified Framework for Zero-Shot Last-Mile Navigation","Mobile manipulation requires a robot to navigate to a target object or receptacle and then execute manipulation. Object-level navigation often stops near the target but fails to produce a manipulation-ready base pose, a challenge addressed as last-mile navigation. UniLM-Nav introduces a zero-shot open-vocabulary framework that decomposes the task into view selection, task-conditioned affordance grounding, and geometry-aware base-pose reasoning, all handled by a shared multimodal large language model backend. Experiments on OVMM show gains over MoTo, and real-world deployment on a Unitree B2 validates applicability.","arXiv :2607 .06537v2 [ cs .RO] 11 Jul 2026  \nUniLM-Nav: A Unified Framework for Zero-Shot  \nLast-Mile Navigation  \nZhuofan Zhang 1 ,2∗, Tianxu Wang2∗, Guoxi Zhang2 , Yixiong Lin2 ,3 , Xilin Wang2 , Hongming Xu2 , Qing Li2 , Song-Chun Zhu 1 ,2 ,4 , Lifeng Fan2†  \n1Tsinghua University  \n2 State Key Laboratory of General Artificial Intelligence, BIGAI  \n3Harbin Institute of Technology  \n4Peking University  \nProject page: [https://unilm-nav.github.io](https://unilm-nav.github.io)  \nAbstract: Mobile manipulation requires a robot to navigate to a target objector receptacle and then perform intended manipulation. However, reaching the vicinity of the target does not guarantee a manipulation-ready base pose, a problem known as last-mile navigation. Prior methods for last-mile navigation either rely on manual pose annotation or task-specific training, limiting their scalability to open-vocabulary settings with fine-grained spatial constraints. We propose UniLM-Nav, a unified framework for zero-shot open-vocabulary last-mile navigation. UniLM-Nav decomposes last-mile navigation into view selection, task-conditioned affordance grounding, and geometry-aware base-pose reasoning, all resolved with a shared multimodal large language model (MLLM) backend. Specifically, UniLM-Nav first selects a reference view that best captures the target object or receptacle from recently collected observations. It then grounds task-relevant affordance point in the selected view and lifts the result into the robot-centric coordinate frame. Finally, conditioned on the grounded affordance, task context, and robot geometry, it infers a manipulation-ready base pose for the robot. We evaluate UniLM-Nav on the OVMM benchmark, where it outperforms the previous state-of-the-art method, MoTo, by 3.13 percentage points. Analyses show that the components of our method are crucial to final performance, and that the choice of MLLM also has a substantial effect. We further deploy UniLM-Navon a Unitree B2 quadruped robot with a 6-DoF Unitree Z1 manipulator, validating its applicability to real-world mobile manipulation tasks.  \nKeywords: Mobile Manipulation, Last-Mile Navigation, MLLM  \n1 Introduction  \nMobile manipulation requires robots to navigate to task-relevant objects or regions and manipulate objects [1, 2], a core capability for household service [3, 4], industrial automation [5], and logistics [6] . In mobile manipulation, the navigation phase should end in a manipulation-ready base pose—a pose that is not only close to the target but also supports the intended manipulation. However, object navigation systems are typically designed to reach the vicinity of the target [7], e.g., within 1–2 meters, which is coarser than manipulation requires. This granularity mismatch can leave the robot poorly aligned or obstructed for manipulation. This motivates the task of last-mile navigation [8]: once near the target, the robot must adjust its base pose to support the manipulation.  \nLast-mile navigation requires jointly identifying interaction-relevant affordance regions and selecting a feasible base pose under spatial and kinematic constraints. Early approaches rely on manual annotation for task-specific base poses [9], limiting their generalization to new tasks or environments.  \n∗Equal contribution. †Corresponding author.  \nTask: Put the bottle on the table, in front of the monitor.  \n1. Object Nav Ends:  \nToo far for manipulation  \n2. Intermediate Pose:  \nManipulation blocked by chair  \n3. Last-Mile Nav Ends:  \nManipulation-ready pose  \nFigure 1: The robot is tasked with placing the bottle on the table in front of the monitor. Object-goal navigation can bring the robot near the table, but the resulting base pose may still be infeasible for placement due to limited reachability or surrounding obstacles. Last-mile navigation refines its base pose based on the task context and robot-scene geometry, thereby enabling feasible placement.  \nAlternatively, learning-based methods","cbCaiiRaabupR9cc","https://ap.wps.com/l/cbCaiiRaabupR9cc","pdf",30197770,2,1,21,"English","en",105,"# Introduction\n## Problem: manipulation-ready base pose and last-mile navigation\n## Proposed method: UniLM-Nav framework\n## Evaluation and real-world deployment","[{\"question\":\"What problem does last-mile navigation solve in mobile manipulation?\",\"answer\":\"It refines the robot’s base pose after it is already near the target so that the pose becomes suitable for the intended manipulation, avoiding cases where object-goal navigation leaves the robot misaligned or blocked.\"},{\"question\":\"How does UniLM-Nav perform zero-shot open-vocabulary last-mile navigation?\",\"answer\":\"UniLM-Nav decomposes the pipeline into view selection, task-conditioned affordance grounding, and geometry-aware base-pose reasoning, using a shared multimodal large language model to complete all phases.\"},{\"question\":\"What evidence shows UniLM-Nav improves over prior work and works in the real world?\",\"answer\":\"On the OVMM benchmark, UniLM-Nav outperforms the prior state of the art MoTo by 3.13 percentage points, and it is deployed on a Unitree B2 quadruped with a 6-DoF Unitree Z1 manipulator for real-world mobile manipulation tasks.\"}]",1784199577,53,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"unilm-nav-a-unified-framework-for-zero-shot-last-mile-navigation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/unilm-nav-a-unified-framework-for-zero-shot-last-mile-navigation/84938/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does last-mile navigation solve in mobile manipulation?","Question",{"text":75,"@type":76},"It refines the robot’s base pose after it is already near the target so that the pose becomes suitable for the intended manipulation, avoiding cases where object-goal navigation leaves the robot misaligned or blocked.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does UniLM-Nav perform zero-shot open-vocabulary last-mile navigation?",{"text":80,"@type":76},"UniLM-Nav decomposes the pipeline into view selection, task-conditioned affordance grounding, and geometry-aware base-pose reasoning, using a shared multimodal large language model to complete all phases.",{"name":82,"@type":73,"acceptedAnswer":83},"What evidence shows UniLM-Nav improves over prior work and works in the real world?",{"text":84,"@type":76},"On the OVMM benchmark, UniLM-Nav outperforms the prior state of the art MoTo by 3.13 percentage points, and it is deployed on a Unitree B2 quadruped with a 6-DoF Unitree Z1 manipulator for real-world mobile manipulation tasks.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]