[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86371-en":3,"doc-seo-86371-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86371,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","SpaceDrive: Infusing Spatial Awareness into VLM-Based Autonomous Driving","End-to-end autonomous driving methods built on vision language models (VLMs) leverage universal visual understanding and strong reasoning from large-scale pretraining, yet struggle with fine-grained 3D spatial relationships needed to interact with the physical world. SpaceDrive introduces spatial information as explicit positional encodings (PEs) rather than textual digit tokens, enabling joint semantic–spatial reasoning. It uses a universal positional encoder over 3D coordinates from multi-view depth, historical egostates, and text prompts, supports 2D token superimposition, and regresses trajectories directly. Experiments show state-of-the-art nuScenes open-loop performance and second-best Bench2Drive closed-loop driving.","SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving  \nPeizheng Li∗ 1 ,2 , Zhenghao Zhang∗ 1 ,4 , David Holtz 1 , Hang Yu 1 ,5 , Yutong Yang 1 ,6 ,  \nYuzhi Lai2 , Rui Song7 , Andreas Geiger2 ,3 , Andreas Zell2  \n1Mercedes-Benz AG, 2University of T¨ubingen, 3T¨ubingen AI Center,  \n4TU Munich, 5 Karlsruhe Institute of Technology, 6University of Stuttgart, 7UCLA  \n[https://zhenghao2519.github.io/SpaceDrive](https://zhenghao2519.github.io/SpaceDrive) Page/  \narXiv :2512 . 10719v3 [ cs .CV] 12 Jul 2026  \nFigure 1 . Spatial awareness in VLM-based end-to-end autonomous driving. (a) Constrained by insufficient 3D pre-training and discrete token-wise encoding, existing end-to-end planners based on the VLM struggle to precisely ground, associate, and predict 3D spatial positions, limiting their planning capabilities. (b) Our proposed SpaceDrive planner introduces a unified 3D coordinate encoding to replace the original VLM’s textual digit tokens and augment visual features, achieving explicit association with 2D perspective semantics to enhance joint spatial reasoning for E2E planning. Compared to current VLM-based methods, it achieves state-of-the-art driving capability in the nuScenes open-loop evaluation and the second-best driving performance in the Bench2Drive closed-loop simulation.  \nAbstract  \nEnd-to-end autonomous driving methods built on vision language models (VLMs) have undergone rapid development driven by their universal visual understanding and strong reasoning capabilities obtained from the large-scale pretraining. However, we find that current VLMs struggle to understand fine-grained 3D spatial relationships which is a fundamental requirement for systems interacting with the physical world. To address this issue, we propose SpaceDrive, a spatial-aware VLM-based driving framework that treats spatial information as explicit positional encodings (PEs) instead of textual digit tokens, enabling joint reasoning over semantic and spatial representations. SpaceDrive employs a universal positional encoder to all 3D coordinates derived from multi-view depth estimation, historical egostates, and text prompts. These 3D PEs are first superimposed to augment the corresponding 2D visual tokens. Meanwhile, they serve as a task-agnostic coordinate rep-  \nresentation, replacing the digit-wise numerical tokens as both inputs and outputs for the VLM. This mechanism enables the model to better index specific visual semantics in spatial reasoning and directly regress trajectory coordinates rather than generating digit-by-digit, thereby enhancing planning accuracy. Extensive experiments validate that SpaceDrive achieves state-of-the-art open-loop performance on the nuScenes dataset and the second-best Driving Score of 78.02 on the Bench2Drive closed-loop benchmark over existing VLM-based methods. Code is available at: [https://github. com/zhenghao2519/SpaceDrive](https://github. com/zhenghao2519/SpaceDrive).  \n1. Introduction  \nLarge-scale pre-trained VLMs are known for their vast knowledge bases and strong reasoning capabilities. Lever-  \n* Equal contribution, names are sorted alphabetically. Correspondence [to:](to: {peizheng.li)[ {](to: {peizheng.li)[peizheng.li](to: {peizheng.li), [zhenghao.zhang](zhenghao.zhang}@mercedes-benz.com)[}](zhenghao.zhang}@mercedes-benz.com)[@mercedes-benz.com](zhenghao.zhang}@mercedes-benz.com).  \naging VLMs to assist [29, 49, 60] or replace [13, 55, 62] traditional end-to-end (E2E) autonomous driving (AD) systems has therefore emerged as a prominent trend recently. These systems typically reformulate AD functions into natural language, and flexibly perform scene understanding, motion prediction and trajectory planning based on semantic information extracted from images. Compared to fixed modular designs [20, 28], VLM-based E2E models promise to achieve superior generalization, addressing increasingly complex and dynamic driving scenarios.  \nHowever, current VLMs demonstrate clear limitatio","cbCaieJdRw2pARcp","https://ap.wps.com/l/cbCaieJdRw2pARcp","pdf",5676800,4,1,18,"English","en",105,"# Introduction\n## Motivation and limitations of current VLM-based 3D reasoning\n## Numerical token modeling issues in language models\n## Proposed SpaceDrive approach and its unified 3D positional encoding","[{\"question\":\"What problem does SpaceDrive target in VLM-based autonomous driving?\",\"answer\":\"Current VLM-based systems have difficulty understanding fine-grained 3D spatial relationships, which are essential for accurate grounding, association, and prediction of 3D positions for planning.\"},{\"question\":\"How does SpaceDrive represent spatial information differently from existing methods?\",\"answer\":\"SpaceDrive treats spatial information as explicit positional encodings derived from 3D coordinates, replacing textual digit tokens and using these encodings as inputs and outputs for the VLM.\"},{\"question\":\"What data sources are used to build the 3D positional encodings in SpaceDrive?\",\"answer\":\"The method derives 3D coordinates from multi-view depth estimation, historical egostates, and text prompts, then superimposes the resulting 3D PEs onto corresponding 2D visual tokens.\"}]",1784211153,45,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"spacedrive-infusing-spatial-awareness-into-vlm-based-autonomous-driving","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/spacedrive-infusing-spatial-awareness-into-vlm-based-autonomous-driving/86371/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does SpaceDrive target in VLM-based autonomous driving?","Question",{"text":75,"@type":76},"Current VLM-based systems have difficulty understanding fine-grained 3D spatial relationships, which are essential for accurate grounding, association, and prediction of 3D positions for planning.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does SpaceDrive represent spatial information differently from existing methods?",{"text":80,"@type":76},"SpaceDrive treats spatial information as explicit positional encodings derived from 3D coordinates, replacing textual digit tokens and using these encodings as inputs and outputs for the VLM.",{"name":82,"@type":73,"acceptedAnswer":83},"What data sources are used to build the 3D positional encodings in SpaceDrive?",{"text":84,"@type":76},"The method derives 3D coordinates from multi-view depth estimation, historical egostates, and text prompts, then superimposes the resulting 3D PEs onto corresponding 2D visual tokens.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]