[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86531-en":3,"doc-seo-86531-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86531,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Learning to Navigate Efficiently with Only 0.58M Trainable Parameters","Learning to Navigate Efficiently with Only 0.58M Trainable Parameters addresses how much model scale a single visual navigation task family truly needs. It proposes a decomposed hybrid navigation architecture where closed-form geometry (e.g., projective relations, occupancy, coordinate transforms) connects three compact learned modules: an egress predictor, a goal-conditioned navigation posterior, and an endpoint-pinned residual diffusion generator. Training uses only 0.58M trainable parameters out of 22.7M total on 44k frames, achieving near state-of-the-art results with low collisions and fast inference.","Learning to Navigate Efficiently with Only 0.58M Trainable  \nParameters  \nEdward Beng Wai Tan∗ , 1 , Siew-Kei Lam 1  \narXiv :2607 . 11029v1 [ cs .RO] 13 Jul 2026  \nAbstract—Recent progress in visual navigation has largely been driven by scale: end-to-end policies with hundreds of millions of parameters trained on billions of frames or large-scale simulated data. We ask how much of this scale a single task family actually requires, and what structure can substitute for it. We propose a decomposed navigation model in which operations with known closed-form structure, such as projective geometry, occupancy, and coordinate transforms, are computed analytically and serve as interfaces between three small learned modules: an egress predictor that grounds the episode goal as a local subgoal in the current view, a navigation predictor that estimates a goal-conditioned posterior over where trajectories travel, and an endpoint-pinned residual diffusion generator that samples trajectory shapes from this posterior. The system trains only 0.58M out of a total of 22.7M parameters, on 44k frames in under one GPU-hour, yet approaches the performance of state-of-the-art models on navigation tasks across 6060 point-goal episodes and 60 environments, while having 233× fewer trainable parameters, the lowest collision rate among all evaluated methods, and 50 Hz inference speed. The decomposition further transfers to no-goal exploration by retraining only the 123k-parameter egress head, and its failure modes under sensor corruption are transparent and analytically correctable.  \nIndex Terms—Vision-Based Navigation; Integrated Planning and Learning; Motion and Path Planning  \nI. INTRODUCTION  \nVisual navigation is a challenging task, requiring robots to navigate to objectives within previously unseen environments. The prevailing route to these capabilities has largely been through scale: reinforcement learning policies trained on billions of steps [1], imitation policies trained on large crossembodiment corpora [2], [3], and more recently, diffusionbased foundation policies trained for trajectory generation on large-scale simulated data [4] . The capability of these models has improved steadily, but at the cost of parameter count, dataset size, and inference cost. These data-driven advancements often bundle many emergent abilities, such as multiple goal types, embodiment awareness, and implicit spatial mapping.  \nHowever, many resource-constrained robots, particularly those which require on-device inference, cannot afford to run these large policies at sufficiently high speed, yet would benefit from the task-specific navigation performance of these models. Motivated by this, we ask the question: for a single task family such as point-goal navigation, how much of this scale is actually required, and specifically what structure facilitates this objective?  \nTo address this, we propose a hybrid approach: combining the strengths of classical methods in performing closed-form  \n1College of Computing and Data Science, Nanyang Technological University, Singapore.  \n∗Corresponding author.  \nFig. 1: Comparison of navigation paradigms. (a) Classical navigation algorithms are efficient but hand-designed heuristics can limit performance,(b) end-to-end navigation models learn geometry, mapping, and control implicitly, at the cost of large parameter and scaling cost,(c) our method retains closedform solutions for operations with known structure, allowing the network to learn the rest of the navigation task with only 0.58M trainable parameters.  \noperations (e.g., projective geometry, occupancy) as the interfaces to reduce the learning load on the neural components, while learning the operations between them. To that end, we factor the point-goal visual navigation task into three distinct operations, each represented by a small, trained operator. The egress predictor learns a local point-goal given the final pointgoal, the navigation predictor learns the goal-cond","cbCait6wwqmVPrbW","https://ap.wps.com/l/cbCait6wwqmVPrbW","pdf",6476846,3,1,6,"English","en",105,"# Introduction\n# Related Work","[{\"question\":\"What core question does the paper ask about visual navigation model scale?\",\"answer\":\"The paper asks how much scale is actually required for a single task family (point-goal navigation) and what structural decomposition can replace brute-force scaling.\"},{\"question\":\"How is the proposed navigation model decomposed?\",\"answer\":\"It factors point-goal navigation into three operations: an egress predictor, a navigation predictor producing a goal-conditioned posterior over trajectory travel, and an endpoint-pinned residual diffusion generator for trajectory shape sampling.\"},{\"question\":\"How many parameters are trained and what performance benefits are reported?\",\"answer\":\"The system trains only 0.58M parameters (from 22.7M total) on 44k frames in under one GPU-hour, reaching near state-of-the-art navigation performance with 233× fewer trainable parameters, low collision rates, and fast 50 Hz inference.\"}]",1784212433,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"learning-to-navigate-efficiently-with-only-058m-trainable-parameters","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/learning-to-navigate-efficiently-with-only-058m-trainable-parameters/86531/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What core question does the paper ask about visual navigation model scale?","Question",{"text":75,"@type":76},"The paper asks how much scale is actually required for a single task family (point-goal navigation) and what structural decomposition can replace brute-force scaling.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is the proposed navigation model decomposed?",{"text":80,"@type":76},"It factors point-goal navigation into three operations: an egress predictor, a navigation predictor producing a goal-conditioned posterior over trajectory travel, and an endpoint-pinned residual diffusion generator for trajectory shape sampling.",{"name":82,"@type":73,"acceptedAnswer":83},"How many parameters are trained and what performance benefits are reported?",{"text":84,"@type":76},"The system trains only 0.58M parameters (from 22.7M total) on 44k frames in under one GPU-hour, reaching near state-of-the-art navigation performance with 233× fewer trainable parameters, low collision rates, and fast 50 Hz inference.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]