[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83368-en":3,"doc-seo-83368-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83368,1099514068365,"Aurelia","https://ap-avatar.wpscdn.com/avatar/10000253d8d9f28188e?_k=1776742907772140068",8,"Research & Report","FSD-VLN Fast-Slow Dual-System Modeling for Aerial Long-Horizon Vision-Language Navigation","Vision-Language Navigation (VLN) supports UAVs by grounding natural-language instructions in real-time visual observations, enabling adaptive autonomy beyond GPS- or pre-programmed approaches. Long-horizon UAV-VLN requires both high-level semantic reasoning and low-latency flight control, yet prior methods face structural inconsistency between global multimodal understanding and temporally coherent action generation. FSD-VLN addresses this via an efficient fast-slow dual-system that decouples semantic priors and low-latency action modeling, improves cross-temporal dependency capture, and uses time-aware adaptive optimization for stable training.","arXiv :2607 .08359v 1 [ cs .RO] 9 Jul 2026  \nFSD-VLN: Fast-Slow Dual-System Modeling for Aerial Long-Horizon Vision-Language Navigation  \nXueke Zhu∗1, Qingyan Meng∗1, Liutao Yu 1 , Wei Zhang 1 , Zhengyu Ma 1 , Huihui Zhou 1 ,2 , Yonghong Tian 1 ,3  \n∗ These authors contributed equally to this work.  \n1 Pengcheng Laboratory  \n2 Shenzhen Institutes of Advanced Technology  \n3 Peking University  \nAbstract  \nVision-Language Navigation (VLN) enables unmanned aerial vehicles (UAVs) to navigate autonomously in unfamiliar environments by grounding natural language instructions in real-time vision observations. Compared to traditional GPS-based or pre-programmed navigation methods, VLN offers more intuitive human-machine interaction and greater adaptability. Effective UAV-VLN systems typically require the integration of high-level semantic reasoning and low-latency flight control. However, existing approaches often suffer from a structural inconsistency between global multimodal understanding and temporally coherent action generation, leading to unstable trajectories and significant decision latency during long-horizon navigation. To address this challenge, we propose FSD-VLN, an efficient fast-slow dualsystem framework that explicitly decouples high-level semantic reasoning from lowlatency action generation. The framework consists of two asynchronous pathways:  \na slow system that extracts stable semantic priors from a pretrained vision–language model, and a fast system built upon a Diffusion Transformer (DiT) that models action distributions for flight command generation. This decoupled design enables stable multi-step prediction and improves trajectory consistency by explicitly capturing cross-temporal action dependencies. Furthermore, a time-aware adaptive optimization strategy is designed to enhance long-horizon training stability and mitigate gradient oscillations during optimization. Extensive experiments in largescale simulated low-altitude environments demonstrate that FSD-VLN achieves up to 2 × higher success rate for navigation in unseen environments compared to prior arts, while reducing both single-action inference latency and overall task execution time by over 50% . These results highlight the importance of explicitly modeling the cooperation between semantic reasoning and temporally coherent control, providing a principled solution for long-horizon aerial VLN.  \nKeywords: Vision-Language Navigation, Long-Horizon Modeling, Fast-Slow Dual-System, Low-Latency Decision Making  \nFSD-VLN  \n Instruction  \n\"Move upwards to a large white building , which is a tall and rectangular structure with balconies . Then , advance forward to it . As you slightly turn right and keep proceeding straight , you will reach it . Slightly turn left and head straight to it . Finally , slightly stop and drop off at it .  \nFast system  \nBVLF: Buffer of vision-language features  \nFigure 1: The navigation pipeline of the proposed FSD-VLN framework. During navigation, the slow system utilizes a pre-trained vision-language model to encode visual observations and language instructions into latent embeddings. The fast system adopts a variant of Diffusion Transformer (DiT), consisting of alternating cross-attention and selfattention modules, to process the UAV state embedding sequence and the visual-language embeddings produced by the slow system, respectively. The fast-slow dual system operates asynchronously to generate low-latency UAV actions.  \n1 Introduction  \nUnmanned Aerial Vehicles (UAVs) are being extensively utilized in a broad spectrum of missions, from environmental monitoring to emergency response, requiring robust decision-making in different scenarios. Traditional UAV navigation often relies heavily on maps or manually designed control strategies[33, 40, 6, 32, 18, 44, 22], making it difficult to adapt to complex, dynamic, and semantically rich task requirements in open environments. Recently, the remarkable breakthroughs in large-scale foun","cbCaimHlXNuZhFe8","https://ap.wps.com/l/cbCaimHlXNuZhFe8","pdf",1442648,5,1,18,"English","en",105,"# Abstract\n# Introduction\n## Background and problem\n## Limitations of existing methods\n# Proposed approach","[{\"question\":\"What problem does FSD-VLN address in long-horizon aerial VLN?\",\"answer\":\"It addresses the structural mismatch between global multimodal semantic understanding and temporally coherent action generation, which causes unstable trajectories and decision latency during long-horizon navigation.\"},{\"question\":\"How does FSD-VLN separate semantic reasoning from action generation?\",\"answer\":\"It uses a slow system to extract stable semantic priors from a pretrained vision-language model and a fast system, based on a Diffusion Transformer variant, to model action distributions for low-latency flight command generation.\"},{\"question\":\"What helps FSD-VLN improve long-horizon training stability and reduce optimization issues?\",\"answer\":\"A time-aware adaptive optimization strategy is used to enhance long-horizon training stability and mitigate gradient oscillations during optimization.\"}]",1784187031,45,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"fsd-vln-fast-slow-dual-system-modeling-for-aerial-long-horizon-vision-language-navigation","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/fsd-vln-fast-slow-dual-system-modeling-for-aerial-long-horizon-vision-language-navigation/83368/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does FSD-VLN address in long-horizon aerial VLN?","Question",{"text":76,"@type":77},"It addresses the structural mismatch between global multimodal semantic understanding and temporally coherent action generation, which causes unstable trajectories and decision latency during long-horizon navigation.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does FSD-VLN separate semantic reasoning from action generation?",{"text":81,"@type":77},"It uses a slow system to extract stable semantic priors from a pretrained vision-language model and a fast system, based on a Diffusion Transformer variant, to model action distributions for low-latency flight command generation.",{"name":83,"@type":74,"acceptedAnswer":84},"What helps FSD-VLN improve long-horizon training stability and reduce optimization issues?",{"text":85,"@type":77},"A time-aware adaptive optimization strategy is used to enhance long-horizon training stability and mitigate gradient oscillations during optimization.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]