[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86021-en":3,"doc-seo-86021-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},86021,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","Traj-VLN Learning Pixel-Space Interaction via Autoregressive Trajectory Generation","Traj-VLN addresses vision-and-language navigation in unseen continuous environments, where an embodied agent must follow natural language instructions through 3D space. Existing vision-language approaches struggle because spatial interaction toward depth directions is hard for VLMs pretrained mainly on 2D RGB conversations. The work proposes fine-tuning VLMs to generate navigation interactions directly in 2D pixel space via autoregressive trajectory generation. Models sequentially predict pixel coordinates and trajectory supervision improves performance, reaching state-of-the-art results with limited compute and data.","Traj-VLN: Learning Pixel-Space Interaction via Autoregressive  \nTrajectory Generation  \nChangfei Fu, Guangcheng Chen, Aoxiang Gu, Haoxiang Liang, Wenjun Xu†, and Hong Zhang†, Life Fellow, IEEE  \narXiv :2607 . 10744v1 [ cs .CV] 12 Jul 2026  \nAbstract—Benefiting from the powerful priors embedded in large-scale pre-training data and the emerging commonsense reasoning ability, large language models (LLMs) have shown unprecedented generalization capabilities in many research fields. Recently, projecting visual embeddings into the language space via vision-language models (VLMs) to achieve sim-toreal and cross-scene generalization has become a prevailing paradigm in the field of Vision-and-Language Navigation in Continuous Environments (VLN-CE). VLN requires an embodied agent to navigate through unseen environments following natural linguistic instructions. We emphasize that a VLN task can be decomposed into a sequence of sub-tasks, each corresponding to a process of 3D spatial interaction with the environments described by instructions such as “walk to the end of the sofa and turn left.” However, such spatial interactions involving moving into the image along the direction of depth sensing are puzzling for VLMs as they were predominantly trained on conversations with RGB images.  \nRather than incorporating depth or 3D geometric information-which VLMs rarely encounter during pretrainingwe propose an alternative approach: fine-tuning VLMs to learn navigation interactions directly in 2D pixel space through autoregressive trajectory generation. Given a linguistic instruction and historical observations, our model sequentially predictsa series of pixel coordinates, drawing a trajectory from the bottom center of the current observation. While prior work has proved that pixel-goal supervision outperforms learning of discrete actions, our experiments further verify that the supervision of pixel-space trajectory significantly enhances VLN performance. Moreover, we demonstrate that our flagship model achieves state-of-the-art level performance with relatively limited computational resources and training data.  \nI. INTRODUCTION  \nVision-and-language navigation (VLN) requires an embodied agent to navigate in unseen environments following natural language instructions, while receiving continuous visual observations from onboard sensors. The instructions are typically composed of detailed sub-task descriptions-for example,“walk out of the dining room, into the living room, and take the first right into the recreation room; stop between the door and the pool table.” Unlike traditional navigation  \nparadigms, which focus on moving the agent from one †Corresponding author ([hzhang@sustech.edu.cn](hzhang@sustech.edu.cn))  \nChangfei Fu and Hong Zhang are with the Shenzhen Key Laboratory of Robotics and Computer Vision, Southern University of Science and Technology, Shenzhen, China. Changfei Fu and Wenjun Xu are also with the Peng Cheng National Laboratory, Shenzhen, China. Weinan Chen is with the State Key Laboratory of Precision Electronic Manufacturing Technology and Equipment, Guangdong University of Technology, Guangzhou, China. This work was supported by the Shenzhen Key Laboratory of Robotics and Computer Vision (ZDSYS20220330160557001), the Major Key Project of PCL (PCL2024A04), and the National Natural Science Foundation of China under Grant U21A20476 .  \nFig. 1: Qualitative comparison of Traj-VLN with general-purpose multi-model large language models (MLLMs) . It’s demonstrated that the popular MLLMs can hardly draw a reasonable trajectory to interact with the 3D environments. With our fine-tuning supervised by pixel-space trajectories, the VLMs achieve the interactions in 2D image plane by learning to autoregressively generate a sequence of pixel coordinates. The prompt used for the general-purpose MLLMs is “please draw a trajectory on the image to follow the instruction: get out of the door and turn left, go through the way between the sof","cbCaimoEu2ARdwIT","https://ap.wps.com/l/cbCaimoEu2ARdwIT","pdf",5450082,6,1,"English","en",105,"# Introduction\n## Task decomposition in VLN\n## Limitations of 2D-trained VLMs for 3D interaction\n## Prior pixel-goal grounding approaches\n## Distribution mismatch and forgetting\n# Proposed Traj-VLN Approach","[{\"question\":\"What problem does Traj-VLN target in vision-and-language navigation?\",\"answer\":\"It targets navigation in unseen continuous environments where an embodied agent must execute 3D spatial interactions described by natural language instructions.\"},{\"question\":\"Why are moving-depth-direction interactions challenging for VLMs?\",\"answer\":\"Because most VLMs are pretrained on conversations with 2D RGB images, so they lack sufficient exposure to 3D information utilization.\"},{\"question\":\"How does Traj-VLN generate navigation trajectories?\",\"answer\":\"Given a linguistic instruction and historical observations, it fine-tunes a VLM to autoregressively predict a sequence of pixel coordinates, forming a trajectory in 2D pixel space.\"}]",1784207860,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"traj-vln-learning-pixel-space-interaction-via-autoregressive-trajectory-generation","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/traj-vln-learning-pixel-space-interaction-via-autoregressive-trajectory-generation/86021/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does Traj-VLN target in vision-and-language navigation?","Question",{"text":75,"@type":76},"It targets navigation in unseen continuous environments where an embodied agent must execute 3D spatial interactions described by natural language instructions.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why are moving-depth-direction interactions challenging for VLMs?",{"text":80,"@type":76},"Because most VLMs are pretrained on conversations with 2D RGB images, so they lack sufficient exposure to 3D information utilization.",{"name":82,"@type":73,"acceptedAnswer":83},"How does Traj-VLN generate navigation trajectories?",{"text":84,"@type":76},"Given a linguistic instruction and historical observations, it fine-tunes a VLM to autoregressively predict a sequence of pixel coordinates, forming a trajectory in 2D pixel space.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":28,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":28,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]