[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83156-en":3,"doc-seo-83156-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},83156,687197207057,"Sage","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","GemNav Discrete-Token Visual Robot Navigation using a Multimodal Large Language Model","GemNav presents a data-efficient visual robot navigation policy that adapts a frozen Multimodal Large Language Model (MLLM) for short-to-medium horizon waypoint navigation using Low-Rank Adaptation (LoRA) only on the language tower. The method avoids a dedicated visual encoder and continuous regression head by using a single discrete token vocabulary shared by waypoints and categorical navigation signals. A soft-decoded auxiliary loss restores metric waypoint structure. Trained on an 8.7-hour open corpus, it transfers zero-shot across four unseen real environments, reaching goals within 0.25–0.42 m in 20 trials, with limited gains from short image histories.","arXiv :2607 .06882v 1 [ cs .RO] 8 Jul 2026  \nGemNav: Discrete-Token Visual Robot Navigation using a Multimodal Large Language Model  \nPeter BohmSaimunur Rahman, Abdelwahed Khamis,  \nSagun Man Singh Shrestha, Chris McCool, Peyman Moghadam CSIRO Technology, Australia  \nAbstract:  \nVisual navigation policies built on large pretrained models have so far followed a common recipe: a dedicated visual encoder, a bespoke action head, and training on thousands of hours of cross-embodiment datasets. We ask whether this recipe is necessary. In this paper, we introduce GemNav, a visual robot navigation policy that adapts a frozen Multimodal Large Language Model (MLLM) for short-to-medium horizon waypoint navigation using Low-Rank Adaptation (LoRA) on the language tower alone, with no auxiliary visual encoder and no continuous regression head. Waypoints and categorical navigation signals share a single discrete token vocabulary generated by the language-model head, and a soft-decoded auxiliary loss recovers the metric structure that pure cross-entropy training discards. On a single 8.7-hour open corpus, roughly three orders of magnitude smaller than competing training sets, the policy transfers zero-shot to four physically distinct unseen environments and stops within 0.25–0.42 m of the goal across 20 real-world trials covering an open carpark, an obstacle carpark, a long outdoor chemical yard, and an indoor warehouse. Conditioning on short image histories improves offline metrics but yields no robot benefit, pointing to a ceiling on what temporal context adds once pretrained vision features are in place. These results indicate that discrete-token adaptation of frozen MLLMs can provide a data-efficient, deployable alternative for foundation model robot navigation.  \nKeywords: Visual Navigation, Data-Efficient Robot Learning, Multimodal Large Language Model (MLLM)  \n1 Introduction  \nRobot navigation has traditionally been built around a tightly coupled stack of perception, state estimation, and control, tuned for a particular robot, sensor setup, or environment [1] . Large pretrained vision-language-action models have shifted this framing, treating navigation as a cross-embodiment learning problem: a single policy trained on hundreds to thousands of hours of trajectories pooled across robots, environments, and sensor configurations [2, 3, 4, 5, 6] .  \nRecent visual navigation policies built on large pretrained models follow a common recipe: a strong visual backbone (often a separate dedicated encoder) is paired with a continuous-action regression head and supervised on a union of multiple navigation datasets collected across embodiments [5, 4, 7] . Yet the scalability of this approach raises a deeper question about where the true source of generalization lies. These policies achieve expansive transferability, but their success has come bundled with an ever-growing need for large demonstration data, specialized action heads, and visual encoders trained from scratch on robotic data. Recent benchmarking further suggests that the visual encoder is the dominant bottleneck for downstream VLA performance [8, 9], pointing toward a trajectory of increasing visual scale as the primary source for improvement. In this paper, we take a different view. The pretraining of modern multimodal large language models [10, 11, 12, 13] has  \n∗ Corresponding author: [peter.bohm@csiro.au](peter.bohm@csiro.au)  \nalready produced visual representations of remarkable breadth and robustness; the question we pose is whether those representations, held frozen, are sufficient for robot navigation, and if the gap between a general-purpose Multimodal Large Language Model (MLLM) and a deployable navigation policy can be closed by rethinking the action representation rather than by scaling perception further.  \nWe ground this question in short-to-medium-horizon waypoint navigation: a higher-level planner supplies the next local goal as a 2D pose, a goal image, or both, and the","cbCaiqHWKaKmZbwR","https://ap.wps.com/l/cbCaiqHWKaKmZbwR","pdf",4505463,1,22,"English","en",105,"# Introduction\n## Waypoint navigation setting\n## GemNav approach\n## Contributions and evaluation","[{\"question\":\"What core idea does GemNav introduce for visual robot navigation?\",\"answer\":\"GemNav adapts a frozen multimodal large language model to waypoint navigation by predicting discrete action tokens, instead of relying on a dedicated visual encoder and a continuous regression head.\"},{\"question\":\"How does GemNav represent waypoints and navigation signals?\",\"answer\":\"Waypoints and categorical decisions (e.g., goal reached/unreachable) share a unified discrete token vocabulary produced by the language-model head in a single token sequence.\"},{\"question\":\"What evidence is provided for transfer and real-world performance?\",\"answer\":\"Trained on a single 8.7-hour open dataset, the policy transfers zero-shot to four unseen, physically distinct environments and stops within 0.25–0.42 m of the goal across 20 real-world trials.\"}]",1784185659,55,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"gemnav-discrete-token-visual-robot-navigation-using-a-multimodal-large-language-model","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/gemnav-discrete-token-visual-robot-navigation-using-a-multimodal-large-language-model/83156/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What core idea does GemNav introduce for visual robot navigation?","Question",{"text":75,"@type":76},"GemNav adapts a frozen multimodal large language model to waypoint navigation by predicting discrete action tokens, instead of relying on a dedicated visual encoder and a continuous regression head.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does GemNav represent waypoints and navigation signals?",{"text":80,"@type":76},"Waypoints and categorical decisions (e.g., goal reached/unreachable) share a unified discrete token vocabulary produced by the language-model head in a single token sequence.",{"name":82,"@type":73,"acceptedAnswer":83},"What evidence is provided for transfer and real-world performance?",{"text":84,"@type":76},"Trained on a single 8.7-hour open dataset, the policy transfers zero-shot to four unseen, physically distinct environments and stops within 0.25–0.42 m of the goal across 20 real-world trials.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]