[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86488-en":3,"doc-seo-86488-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86488,7971461740886,"Theodore","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","ABot-N1: Toward a General Visual Language Navigation Foundation Model","Visual language navigation foundation models unify deep reasoning for spatial decisions with broad versatility across embodied tasks, yet current monolithic policies often suffer from coordinate drift, limited long-tail semantics, and black-box behavior that reduces interpretability. ABot-N1 addresses these issues by decoupling cognition from control using a slow-fast dual visual-language system. A slow reasoner performs explicit chain-of-thought and outputs pixel goal anchors as a universal interface, while a fast expert generates continuous waypoints from textual cues and pixel guidance. The model delivers state-of-the-art urban-scale results, improving POI arrival to 77.3% and achieving high SR in complex indoor and outdoor scenes.","arXiv :2607 . 10383v1 [ cs .CV] 11 Jul 2026  \nABot-N1: Toward a General Visual Language Navigation  \nFoundation Model  \nAbstract  \nVisual Language Navigation foundation models aim to unify deep reasoning for grounded spatial decisions with broad versatility for diverse embodied tasks. Current approaches typically achieve this integration via monolithic policies that map observations directly to actions, yet they often suffer from coordinate drift and poor handling of long-tail semantics. Furthermore, these black-box mappings lack interpretability, hindering the simultaneous achievement of generality, robustness, and transparency. We present ABot-N1, a step toward a general Visual Language Navigation foundation model, that addresses these challenges by decoupling cognition from control via aslow-fast architecture guided by dual visual-language signals. More specifically, a slow visionlanguage reasoner performs explicit Chain-of-Thought reasoning while producing a pixel goal. This compact set of image-space anchor points serves as a universal interface for diverse tasks, including point-goal, object-goal, poi-goal, instruction-following, and person-following. Subsequently, a fast action expert leverages both the textual cues and the pixel guidance to generate continuous waypoints at the native control frequency. By bridging high-level intents and low-level control through pixel-grounded anchors paired with explicit linguistic traces, our approach ensures robust, generalizable, and interpretable navigation across simulation and real-world benchmarks. ABotN1 establishes new state-of-the-art records, delivering massive gains specifically in urban-scale navigation: boosting POI arrival by 35.0%(to 77.3%) and achieving 95.4%/92.9% SR in complex indoor and outdoor scenes. It also maintains superior robustness across object-reaching, person-following, and instruction-following tasks. New Point-Goal/POI-Goal benchmarks are released as open source to advance the field of urban-scale navigation.  \nProject Page: [https://amap-cvlab.github.io/ABot-Navigation/ABot-N1/](https://amap-cvlab.github.io/ABot-Navigation/ABot-N1/)  \nContents  \n1 Introduction ................................................ 3  \n2 Related Works .............................................. 5  \n2.1 Generalist Navigation Foundation Models ............................. 6  \n2.2 Brain-Body Decoupling: Dual-System VLN Architectures .................... 6  \n2.3 Toward General Embodied Reasoning for Navigation ....................... 6  \n3 Preliminaries ............................................... 7  \n3.1 Embodied Navigation As Goal-Conditioned Visual Control ................... 7  \n3.2 Unifying FIVE Navigation Tasks in ONE Framework ....................... 7  \n3.3 VLM-Conditioned Policy and Continuous-Action Decoding ................... 8  \n4 Methods .................................................. 8  \n4.1 Model Architecture .......................................... 8  \n4.2 Pretraining .............................................. 10  \n4.2.1 Training Protocol ....................................... 10  \n4.2.2 Pretraining Data Recipe ................................... 11  \n4.3 Post-Training ............................................. 16  \n4.3.1 Formulation .......................................... 16  \n4.3.2 Reward Design ........................................ 16  \n4.3.3 Balance Sampling for GRPO Training Dataset ...................... 17  \n4.3.4 Training Dynamics and Generality ............................. 19  \n5 Benchmark ................................................ 19  \n5.1 ABotN-PointBench .......................................... 19  \n5.2 ABotN-POIBench .......................................... 20  \n6 Experiments ................................................ 21  \n6.1 Simulation Evaluation ........................................ 21  \n6.1.1 Instruction-Following: VLN-CE R2R/RxR ........................ 22  \n6.1.2 Object-Goal: Short-Horizon OVON ...","cbCaigDZYmXrPGJV","https://ap.wps.com/l/cbCaigDZYmXrPGJV","pdf",12457178,3,1,36,"English","en",105,"# Introduction\n# Related Works\n## Generalist Navigation Foundation Models\n## Brain-Body Decoupling: Dual-System VLN Architectures\n## Toward General Embodied Reasoning for Navigation\n# Preliminaries\n## Embodied Navigation As Goal-Conditioned Visual Control\n## Unifying FIVE Navigation Tasks in ONE Framework\n## VLM-Conditioned Policy and Continuous-Action Decoding\n# Methods\n## Model Architecture\n## Pretraining\n## Post-Training\n# Benchmark\n## ABotN-PointBench\n## ABotN-POIBench\n# Experiments\n## Simulation Evaluation\n## Real-World Deployment\n# Conclusion\n# Contributions","[{\"question\":\"What problem does ABot-N1 target in visual language navigation foundation models?\",\"answer\":\"It targets coordinate drift, weak long-tail semantic handling, and poor interpretability caused by monolithic black-box observation-to-action policies, which limit generality, robustness, and transparency.\"},{\"question\":\"How does ABot-N1 decouple cognition and control?\",\"answer\":\"ABot-N1 uses a slow-fast dual visual-language system: a slow vision-language reasoner produces explicit reasoning and pixel goal anchors, and a fast action expert converts textual cues plus pixel guidance into continuous waypoints at control frequency.\"},{\"question\":\"What performance gains does ABot-N1 report for urban-scale navigation?\",\"answer\":\"It boosts POI arrival by 35.0% to 77.3% and reports 95.4%/92.9% SR for complex indoor and outdoor scenes, while maintaining strong robustness across object-reaching, person-following, and instruction-following.\"}]",1784212107,91,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"abot-n1-toward-a-general-visual-language-navigation-foundation-model","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/abot-n1-toward-a-general-visual-language-navigation-foundation-model/86488/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does ABot-N1 target in visual language navigation foundation models?","Question",{"text":75,"@type":76},"It targets coordinate drift, weak long-tail semantic handling, and poor interpretability caused by monolithic black-box observation-to-action policies, which limit generality, robustness, and transparency.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does ABot-N1 decouple cognition and control?",{"text":80,"@type":76},"ABot-N1 uses a slow-fast dual visual-language system: a slow vision-language reasoner produces explicit reasoning and pixel goal anchors, and a fast action expert converts textual cues plus pixel guidance into continuous waypoints at control frequency.",{"name":82,"@type":73,"acceptedAnswer":83},"What performance gains does ABot-N1 report for urban-scale navigation?",{"text":84,"@type":76},"It boosts POI arrival by 35.0% to 77.3% and reports 95.4%/92.9% SR for complex indoor and outdoor scenes, while maintaining strong robustness across object-reaching, person-following, and instruction-following.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]