[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83369-en":3,"doc-seo-83369-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83369,1099514068365,"Aurelia","https://ap-avatar.wpscdn.com/avatar/10000253d8d9f28188e?_k=1776742907772140068",8,"Research & Report","WCog-VLA: A Dual-Level World-Cognitive Vision-Language-Action Model for End-to-End Autonomous Driving","Vision-Language-Action (VLA) models have advanced end-to-end autonomous driving, but existing approaches often lack comprehensive world cognition or suffer fragmented world foresight, limiting behavior to reactive driving. WCog-VLA introduces a dual-level World-Cognitive VLA framework that unifies semantic world forecasting with generative world evolution for proactive driving. At the semantic level, it integrates 3D spatial perception, agent tokens, and Game-CoT reasoning; at the generative level, ADDT synthesizes physically-plausible multi-agent trajectories with fewer denoising steps. Experiments on NAVSIM report a SOTA PDMS of 92.9.","arXiv :2607 .08375v 1 [ cs .CV] 9 Jul 2026  \nWCog-VLA: A Dual-Level World-Cognitive Vision-Language-Action Model for End-to-End Autonomous Driving  \nXuerun Yan 1 ,2†, Zhexi Lian 1†, Nuoheng Zhang 1 , Shiyu Fang 1 , Haoran Wang 1 , Chen Lv2 , Jia Hu 1⊠, and Binyang Song2⊠  \n1 Tongji University, China  \n2 Nanyang Technological University, Singapore  \n† Equal contribution ⊠ Corresponding author  \nAbstract. Vision-Language-Action (VLA) models have advanced endto-end autonomous driving. However, existing methods either lack comprehensive world cognition or suffer from fragmented world foresight, inherently confining these models to reactive driving. To address this limitation, we propose WCog-VLA, a novel dual-level World-Cognitive VLA framework that successfully bridges semantic world forecasting with generative world evolution to achieve proactive autonomous driving.  \nAt the semantic level, WCog-VLA unifies world cognition and reasoning by incorporating 3D spatial perception and injecting agent tokens to capture the world dynamics, while concurrently enabling Game-theoretic Chain-of-Thought (Game-CoT) reasoning. At the generative level, we introduce the Aligned Decoupled Diffusion Transformer (ADDT) as a powerful generative world model that synthesizes physically-plausible joint multi-agent trajectories. Through scene representation alignment, ADDT reduces the number of denoising steps required and thus significantly accelerates inference. To facilitate strategic reasoning, we further construct a large-scale dataset featuring 85k Game-CoT annotations.  \nExtensive experiments on the NAVSIM benchmark demonstrate that WCog-VLA achieves a State-Of-The-Art (SOTA) PDMS score of 92.9 .  \nKeywords: End-to-end autonomous driving · Vision-Language-Action  \n· World cognition  \n1 Introduction  \nEnd-to-end (E2E) autonomous driving has emerged as a dominant paradigm by directly mapping raw sensory inputs to planned trajectories [5, 19, 20, 26] within a unified and differentiable framework. Although these E2E models show remarkable performance in common scenarios, they often struggle in complex or long-tail situations [4, 63] . This fragility arises from insufficient causal reasoning and world knowledge, leaving the models unable to fully understand and reason about the surrounding environments.  \n2 X. Yan et al.  \nFig. 1: Four paradigms of leveraging VLM in E2E autonomous driving. Our method (d) advances existing frameworks to enable proactive driving by establishing a dual-level world cognition with the integration of semantic forecasting and generative evolution.  \nTo address these long-tail challenges, Vision-Language-Models (VLMs) [1, 2] have been increasingly integrated into E2E autonomous driving frameworks [25, 47, 53] . Equipped with extensive world knowledge and strong reasoning capabilities, VLMs significantly advance scene comprehension in complex driving scenarios [27,60]. Building upon this foundation, Vision-Language-Action (VLA) models further extend VLM capabilities to action generation. Existing VLA approaches typically operate in two ways. The first formulates action outputs as autoregressive sequence generation (Fig. 1(a)), producing either discrete text tokens [37, 39, 42] or learned action codes [64] . Alternatively, the second utilizes VLMs as cognitive encoders and attaches dedicated action decoders (Fig. 1(b)) to generate continuous trajectories [13, 52], e.g ., diffusion models [23, 32] .  \nDespite these advancements, existing VLA models still face several challenges: (1) Lack of 3D spatial awareness. Relying primarily on 2D image features, current models lack structured 3D spatial representations of surrounding road participants [38, 63], which are essential for accurate spatial reasoning and precise ego planning. (2) Insufficient world cognition. Existing methods struggle to adequately represent world states and forecast future dynamics [46], such as the intentions of surrounding agents. This confines current method","cbCaifvB7m8JSgzs","https://ap.wps.com/l/cbCaifvB7m8JSgzs","pdf",2584195,3,1,20,"English","en",105,"# Introduction\n## Background and Motivation\n## Challenges in Existing VLA Methods\n## Proposed WCog-VLA Approach","[{\"question\":\"What problem does WCog-VLA address in existing VLA-based autonomous driving?\",\"answer\":\"It addresses the lack of comprehensive world cognition and fragmented world foresight that confines models to reactive driving, especially in complex long-tail scenarios.\"},{\"question\":\"How does WCog-VLA enable proactive driving rather than reactive behavior?\",\"answer\":\"It bridges semantic world forecasting and generative world evolution, using semantic-level Game-CoT reasoning plus a generative world model to synthesize joint interactive trajectories.\"},{\"question\":\"What role does ADDT play in WCog-VLA?\",\"answer\":\"ADDT is the generative world model that generates physically-plausible joint multi-agent trajectories, and scene representation alignment accelerates inference by reducing denoising steps.\"}]",1784187040,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"wcog-vla-a-dual-level-world-cognitive-vision-language-action-model-for-end-to-end-autonomous-driving","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/wcog-vla-a-dual-level-world-cognitive-vision-language-action-model-for-end-to-end-autonomous-driving/83369/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does WCog-VLA address in existing VLA-based autonomous driving?","Question",{"text":75,"@type":76},"It addresses the lack of comprehensive world cognition and fragmented world foresight that confines models to reactive driving, especially in complex long-tail scenarios.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does WCog-VLA enable proactive driving rather than reactive behavior?",{"text":80,"@type":76},"It bridges semantic world forecasting and generative world evolution, using semantic-level Game-CoT reasoning plus a generative world model to synthesize joint interactive trajectories.",{"name":82,"@type":73,"acceptedAnswer":83},"What role does ADDT play in WCog-VLA?",{"text":84,"@type":76},"ADDT is the generative world model that generates physically-plausible joint multi-agent trajectories, and scene representation alignment accelerates inference by reducing denoising steps.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":29,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":22,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":22,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]