[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86256-en":3,"doc-seo-86256-105":30,"detail-sidebar-cat-0-en-105":84},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86256,687197207919,"Theodora","https://ap-avatar.wpscdn.com/avatar/a000253d6f5f7c60be?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779446848396160552",8,"Research & Report","Parse Search and Confirmation Training Free Aerial Vision and Dialog Navigation with Chain of Thought Reasoning and Structured Spatial Memory","This paper addresses the Aerial Vision-and-Dialog Navigation (AVDN) task in a training-free setting for resource-efficient high-altitude UAV navigation. Directly applying multimodal large language models can yield unreliable trajectories due to weak directional grounding and missing explicit spatial memory. PSC-AVDN introduces a three-stage Parsing-Search-Confirmation pipeline tightly coupled with Structured Spatial Memory (SSM). Parsing produces stable geometric cues, Search uses chain-of-thought exploration, and Confirmation verifies candidate regions while SSM fuses multi-scale observations, visual memory, and structured geometry for long-horizon consistency.","Parse, Search, and Confirmation: Training-Free Aerial Vision-and-Dialog Navigation with Chain-of-Thought Reasoning and Structured Spatial Memory  \nYu Qi 1 Hongyu Li4 Shaofei Huang5 Tianrui Hui 1 ,2 ,3 *  \nYaxiong Wang 1 Lechao Cheng 1 Zhun Zhong 1 * Si Liu4 Meng Wang 1  \n1 School of Computer Science and Information Engineering, Hefei University of Technology  \n2Jianghuai Advance Technology Center 3Anhui Provincial Key Laboratory of Humanoid Robots  \n4 School of Artificial Intelligence, Beihang University 5University of Macau  \narXiv :2607 . 11529v1 [ cs .CV] 13 Jul 2026  \nAbstract  \nIn this paper, we tackle the Aerial Vision-and-Dialog Navigation (AVDN) task in the training-free setting for resourceefficient high-altitude UAV navigation.Naively applying MLLMs leads to unreliable navigation due to weak directional grounding and the lack of explicit spatial memory. To address these issues, we propose PSC-AVDN, a trainingfree framework that tightly couples a three-stage ParsingSearch-Confirmation reasoning pipeline with a Structured Spatial Memory (SSM) . The parsing stage uses an LLM to convert ambiguous dialogue instructions into stable geometric directional and destination cues. A Search Chainof-Thought (S-CoT) then performs stepwise target exploration under high-altitude observations, and a Confirmation Chain-of-Thought (C-CoT) conducts fine-grained verification around candidate regions to resolve visual ambiguity. Meanwhile, SSM integrates three complementary sources of spatial cues, including multi-scale visual observation, spatial visual memory, and structured geometric memory to provide global spatial context and long-horizon consistency. Extensive experiments on ANDH and ANDHFull show that PSC-AVDN establishes new state-of-the-art performance in the training-free setting, matching or surpassing several finetuned methods.Code will be publicly available at: [https://github.com/QY6616/PSC-AVDN](https://github.com/QY6616/PSC-AVDN)  \n1. Introduction  \nAerial Vision-and-Dialogue Navigation (AVDN) [7] enables UAVs to follow language instructions while resolving ambiguities through questions and interpreting contextual cues along their flight trajectory. AVDN adopts a highaltitude, top-down perspective similar to remote sensing imagery. Compared with low-altitude urban settings [14, 19,  \n*Corresponding authors.  \nFigure 1 . Motivation of our method. (a) The MLLM baseline suffers from ambiguous directional descriptions and the domain gap between high-altitude imagery and ground-level training data, leading to inaccurate localization. (b) Our PSC-AVDN eliminates directional ambiguity through instruction parsing, performs structured search via chain-of-thought reasoning, and conducts finegrained confirmation around the candidate region to achieve more reliable navigation. In addition, a structured spatial memory is introduced to provide clearer spatial context for reasoning.  \n37, 40, 45], this high-altitude setting covers a broader visual scope where landmarks are small in scale and can be either densely clustered or sparsely distributed, making it challenging to localize and track fine-grained landmarks during navigation. These characteristics make AVDN well-suited for applications requiring large-scale environmental understanding, such as disaster rescue, environmental monitoring, and geospatial mapping [18, 29, 31, 38], where operators must reason over wide areas relying on visually subtle landmarks.  \nIn practice, however, AVDN methods rely on supervised finetuning [7, 35, 36], which entails high computational costs, heavy annotation efforts, and repeated re-  \nannotation and re-training when adapting to new environments. To overcome these issues, we explore a trainingfree AVDN framework that enables resource-efficient UAV navigation from high-altitude perspectives, which is particularly difficult without task-specific training. Thanks to recent advances in Multimodal Large Language Models (MLLMs) [2, 11, 21, 26, 47], we cons","cbCaiigPeUGhAGnN","https://ap.wps.com/l/cbCaiigPeUGhAGnN","pdf",2800504,6,1,10,"English","en",105,"# Abstract\n# Introduction\n## Motivation and challenges of high-altitude AVDN\n## Training-free approach with MLLM baseline and its limitations\n## PSC-AVDN: Parsing-Search-Confirmation and Structured Spatial Memory","[{\"question\":\"How does PSC-AVDN improve reliability compared with the baseline?\",\"answer\":\"PSC-AVDN separates directional understanding from target localization via a Parsing-Search-Confirmation pipeline and adds Structured Spatial Memory that combines multi-scale visual observations, spatial visual memory, and structured geometric memory for better global context.\"}]",1784209853,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":79,"head_meta":81,"extra_data":83,"updated_unix":28},"parse-search-and-confirmation-training-free-aerial-vision-and-dialog-navigation-with-chain-of-thought-reasoning-and-structured-spatial-memory","",{"@graph":36,"@context":78},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/parse-search-and-confirmation-training-free-aerial-vision-and-dialog-navigation-with-chain-of-thought-reasoning-and-structured-spatial-memory/86256/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72],{"name":73,"@type":74,"acceptedAnswer":75},"How does PSC-AVDN improve reliability compared with the baseline?","Question",{"text":76,"@type":77},"PSC-AVDN separates directional understanding from target localization via a Parsing-Search-Confirmation pipeline and adds Structured Spatial Memory that combines multi-scale visual observations, spatial visual memory, and structured geometric memory for better global context.","Answer","https://schema.org",{"og:url":52,"og:type":80,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":82,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":85},[86,90,94,98,103,107,112,115,120,123,126],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":87,"show_sort_weight":88,"slug":89},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":91,"show_sort_weight":92,"slug":93},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Exam",70,"exam",{"id":99,"doc_module":4,"doc_module_name":46,"category_name":100,"show_sort_weight":101,"slug":102},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":22,"slug":125},"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":99,"slug":129},19,"General","general"]