[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"detail-sidebar-cat-0-en-105":3,"doc-seo-327872-105":59,"doc-detail-327872-en":130},{"code":4,"msg":5,"data":6},0,"success",[7,13,18,23,28,33,38,43,48,51,55],{"id":8,"doc_module":4,"doc_module_name":9,"category_name":10,"show_sort_weight":11,"slug":12},1,"Document","Story & Novel",90,"story-novel",{"id":14,"doc_module":4,"doc_module_name":9,"category_name":15,"show_sort_weight":16,"slug":17},2,"Literature",80,"literature",{"id":19,"doc_module":4,"doc_module_name":9,"category_name":20,"show_sort_weight":21,"slug":22},4,"Exam",70,"exam",{"id":24,"doc_module":4,"doc_module_name":9,"category_name":25,"show_sort_weight":26,"slug":27},5,"Comic",60,"comic",{"id":29,"doc_module":4,"doc_module_name":9,"category_name":30,"show_sort_weight":31,"slug":32},6,"Technology",50,"technology",{"id":34,"doc_module":4,"doc_module_name":9,"category_name":35,"show_sort_weight":36,"slug":37},7,"Healthcare",40,"healthcare",{"id":39,"doc_module":4,"doc_module_name":9,"category_name":40,"show_sort_weight":41,"slug":42},8,"Research & Report",30,"research-report",{"id":44,"doc_module":4,"doc_module_name":9,"category_name":45,"show_sort_weight":46,"slug":47},9,"Religion & Spirituality",20,"religion-spirituality",{"id":46,"doc_module":4,"doc_module_name":9,"category_name":49,"show_sort_weight":46,"slug":50},"World Cup","world-cup",{"id":52,"doc_module":4,"doc_module_name":9,"category_name":53,"show_sort_weight":52,"slug":54},10,"Lifestyle","lifestyle",{"id":56,"doc_module":4,"doc_module_name":9,"category_name":57,"show_sort_weight":24,"slug":58},19,"General","general",{"code":4,"msg":60,"data":61},"ok",{"site_id":62,"language":63,"slug":64,"title":65,"keywords":66,"description":67,"schema_data":68,"social_meta":123,"head_meta":125,"extra_data":127,"updated_unix":129},105,"en","critiquedrivevlm-from-verifier-guided-reinforcement-learning-to-latent-thought-distillation-for-autonomous-driving","CritiqueDriveVLM - From Verifier-Guided Reinforcement Learning to Latent Thought Distillation for Autonomous Driving","","End-to-end Vision-Language Models (VLMs) show strong promise for autonomous driving, yet standard Supervised Fine-Tuning (SFT) often yields reasoning hallucinations and conservative biases. Tool-augmented and Chain-of-Thought (CoT) methods reduce some issues but introduce excessive token usage and unacceptable latency, blocking real-time deployment. CritiqueDriveVLM resolves this reliability-efficiency trade-off through a unified three-stage framework internalizing reasoning within the VLM.",{"@graph":69,"@context":122},[70,84,105],{"@type":71,"itemListElement":72},"BreadcrumbList",[73,77,79,82],{"item":74,"name":75,"@type":76,"position":8},"https://docshare.wps.com","Home","ListItem",{"item":78,"name":9,"@type":76,"position":14},"https://docshare.wps.com/document/",{"item":80,"name":40,"@type":76,"position":81},"https://docshare.wps.com/document/research-report/",3,{"item":83,"name":65,"@type":76,"position":19},"https://docshare.wps.com/document/critiquedrivevlm-from-verifier-guided-reinforcement-learning-to-latent-thought-distillation-for-autonomous-driving/327872/",{"url":83,"name":65,"@type":85,"image":86,"author":91,"headline":65,"publisher":94,"fileFormat":97,"inLanguage":63,"description":67,"dateModified":98,"datePublished":99,"encodingFormat":97,"isAccessibleForFree":100,"interactionStatistic":101},"DigitalDocument",{"url":87,"@type":88,"width":89,"height":90},"https://docshare.wps.com/thumbnails/critiquedrivevlm-from-verifier-guided-reinforcement-learning-to-latent-thought-distillation-for-autonomous-driving/327872.png","ImageObject",300,407,{"name":92,"@type":93},"Levi","Person",{"url":74,"name":95,"@type":96},"DocShare","Organization","application/pdf","2026-09-22","2026-09-21",true,{"@type":102,"interactionType":103,"userInteractionCount":8},"InteractionCounter",{"@type":104},"ViewAction",{"@type":106,"mainEntity":107},"FAQPage",[108,114,118],{"name":109,"@type":110,"acceptedAnswer":111},"Why does standard Supervised Fine-Tuning (SFT) underperform in autonomous driving with VLMs?","Question",{"text":112,"@type":113},"Standard SFT tends to mimic surface-level trajectory distributions without truly internalizing deep logical deduction, leading to reasoning hallucinations and conservative biases in complex dynamic scenes.","Answer",{"name":115,"@type":110,"acceptedAnswer":116},"What are the drawbacks of tool-augmented frameworks and Chain-of-Thought (CoT) approaches?",{"text":117,"@type":113},"They rely on brittle external APIs and generate many explicit reasoning tokens, which increases computational overhead and latency, making real-time deployment impractical.",{"name":119,"@type":110,"acceptedAnswer":120},"How does CritiqueDriveVLM improve both reliability and efficiency?",{"text":121,"@type":113},"It uses Critique-Driven Multi-Turn Reinforcement Learning with a multi-dimensional verifier to internalize logical deduction, then applies Latent Thought Distillation to compress the teacher’s converged reasoning into a fast, CoT-free student.","https://schema.org",{"og:url":83,"og:type":124,"og:title":65,"og:site_name":95,"og:description":67},"article",{"robots":126,"canonical":83},"index,follow",{"doc_id":128,"site_id":62},327872,1790084033,{"code":4,"msg":5,"data":131},{"doc_id":128,"user_id":132,"nickname":92,"user_avatar":133,"doc_module":4,"category_id":39,"category_name":40,"doc_title":65,"doc_description":67,"doc_content":134,"file_id":135,"file_url":136,"file_type":137,"file_size":138,"view_count":8,"is_deleted":4,"is_public":8,"is_downloadable":8,"audit_status":8,"page_count":139,"language":140,"language_code":63,"site_id":62,"html_lang":63,"table_of_contents":141,"faqs":142,"seo_title":143,"seo_description":67,"update_tm":144,"read_time":145},7971461740909,"https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d","arXiv :2607 .04 179v 1 [ cs .CV] 5 Jul 2026  \nCritiqueDriveVLM: From Verifier-Guided Reinforcement Learning to Latent Thought Distillation for Autonomous Driving  \nZhaohong Liu 1 ,3 , Hao Ye 1 ,3 , Xianlin Zhang2 ,3 , and Mengshi Qi 1 ,3 (B)  \n1 State Key Laboratory of Networking and Switching Technology  \n2 School of Digital Media & Design Arts  \n3 Beijing University of Posts and Telecommunications, Beijing, China  \n{liuzhaoh, haoye, zxlin, [qms}@bupt.edu.cn](qms}@bupt.edu.cn)  \nAbstract. End-to-end Vision-Language Models (VLMs) show immense potential in autonomous driving. However, standard Supervised FineTuning (SFT) often suffers from reasoning hallucinations and conservative biases. While traditional tool-augmented frameworks and Chain-ofThought (CoT) approaches mitigate these issues, they incur exorbitant token consumption and unacceptable latency, rendering real-time deployment impractical. To resolve this reliability-efficiency trade-off, we propose CritiqueDriveVLM, a novel unified three-stage framework internalizing reasoning directly into the VLM. First, we introduce Critique-Driven Multi-Turn Reinforcement Learning (RL) guided by a multi-dimensional verifier. By providing granular scalar feedback and a multi-turn penalty, we force the policy to internalize logical deduction, cultivating a robust System-2 Teacher that achieves high accuracy without fragile external tools. Subsequently, we propose Latent Thought Distillation to overcome the latency bottleneck. By aligning the Student’s latent representations with the Teacher’s fully converged reasoning states, we compress deep logical capabilities into a fast, CoT-free System-1 Student. Extensive experiments on the widely-used DriveLMM-o1 benchmark demonstrate remarkable improvements. Compared to the base model, our tool-free Teacher significantly boosts Multiple Choice Quality (MCQ) from 55.54% to a state-of-the-art 76 .54% . Crucially, our distilled Student preserves competitive reasoning depth while drastically minimizing generation length to an average of merely 28 tokens. This slashes inference latency by 88%(from 3482 ms to 416 ms), paving a highly robust pathway for low-latency autonomous driving. Our source code is available at [https://github.com/MICLAB-BUPT/CritiqueDriveVLM](https://github.com/MICLAB-BUPT/CritiqueDriveVLM).  \nKeywords: Vision-Language Models · Autonomous Driving · Reinforcement Learning · Knowledge Distillation  \n1 Introduction  \nThe integration of Vision-Language Models (VLMs) [1, 2, 4, 28, 43] has catalyzed a profound paradigm shift in autonomous driving, transitioning the field from  \n2 Z. Liu et al.  \nQuestion: What behavior is least likely to cause a traffic incident? Choose from the following answers only.  \nA) Slow down. B) Move right. C) Speed up.  \nD) Maintain speed. E) Move left. F) None of the options.  \nFig. 1: Paradigm comparison of VLM-based autonomous driving. (a) Standard SFT is prone to reasoning hallucinations and conservative biases. (b) Tool-Augmented methods suffer from brittle external APIs and high latency. (c) Critique-Driven RL (Teacher) internalizes deep logic without relying on external tools. (d) Latent Thought Distillation (Student) enables instant, CoT-free execution, eliminating the overhead of explicit reasoning tokens.  \ntraditional, modular perception-planning pipelines [8,9,23,25] toward end-to-end holistic reasoning frameworks [16,19] . By leveraging extensive world knowledge and powerful visual-semantic alignment, VLMs demonstrate remarkable potential in interpreting complex traffic scenes and understanding nuanced dynamic interactions. However, deploying these models directly into safety-critical environments remains an open challenge. Current methodologies predominantly rely on standard Supervised Fine-Tuning (SFT) [47] to align pre-trained VLMs with driving-specific instructions. Unfortunately, standard SFT merely mimics the surface-level statistical distribution of human driving trajectories without ","cbCaiiCWHyWXfiWl","https://ap.wps.com/l/cbCaiiCWHyWXfiWl","pdf",1728303,18,"English","# Abstract\n# Keywords\n# 1 Introduction\n## SFT limitations and reasoning failures\n## Tool-augmented and CoT approaches: fragility and latency\n## Critique-Driven Multi-Turn RL and verifier guidance\n## Latent Thought Distillation for low-latency execution","[{\"question\":\"Why does standard Supervised Fine-Tuning (SFT) underperform in autonomous driving with VLMs?\",\"answer\":\"Standard SFT tends to mimic surface-level trajectory distributions without truly internalizing deep logical deduction, leading to reasoning hallucinations and conservative biases in complex dynamic scenes.\"},{\"question\":\"What are the drawbacks of tool-augmented frameworks and Chain-of-Thought (CoT) approaches?\",\"answer\":\"They rely on brittle external APIs and generate many explicit reasoning tokens, which increases computational overhead and latency, making real-time deployment impractical.\"},{\"question\":\"How does CritiqueDriveVLM improve both reliability and efficiency?\",\"answer\":\"It uses Critique-Driven Multi-Turn Reinforcement Learning with a multi-dimensional verifier to internalize logical deduction, then applies Latent Thought Distillation to compress the teacher’s converged reasoning into a fast, CoT-free student.\"}]","CritiqueDriveVLM - From Verifier-Guided Reinforcement Learning to Latent Thought Distillation for Autonomous Driving | PDF",1789978234,45]