[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83482-en":3,"doc-seo-83482-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},83482,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1782698725881665579",8,"Research & Report","DriveVer: Lightweight Trajectory Evaluator as Test-Time Verifier for Autonomous Driving","End-to-end autonomous driving systems can suffer performance bottlenecks because training-time scaling increases compute cost while providing diminishing returns, and many planners generate a single trajectory without any inference-time validation. DriveVer addresses this gap with a lightweight, plug-and-play test-time verifier that evaluates candidate trajectories and refines them safely without heavy retraining. A NAVSIM-based dataset is built via condition-driven clustering and balanced sampling. A dual-head model predicts safety confidence and absolute geometric refinement, improving base planners with minimal overhead and real-time efficiency.","DriveVer: Lightweight Trajectory Evaluator as Test-Time Verifier for  \nAutonomous Driving  \nChong He 1* , Yuechen Luo 1* , Fang Li2 , Shaoqing Xu2 , Fuxi Wen 1,✉  \narXiv :2607 .00399v 1 [ cs .CV] 1 Jul 2026  \nAbstract—End-to-end autonomous driving models often encounter performance bottlenecks, as training-time scaling leads to high computational costs and diminishing marginal returns. Existing planners typically adopt a one-shot generation paradigm, lacking secondary validation and active correction mechanisms to detect and revise suboptimal or unsafe trajectories during inference. To address this issue, we propose DriveVer, a lightweight, plug-and-play Test-Time Verifier that leverages the test-time scaling paradigm to enable autonomous driving systems to validate and refine trajectories without costly and heavy training. We construct a dedicated trajectory dataset based on the NAVSIM benchmark through condition-driven clustering and balanced sampling according to ego-vehicle states and navigation commands. Employing a dual-head architecture, DriveVer efficiently fuses candidate trajectories with multi-view visual representations and ego-vehicle kinematic features to simultaneously predict a safety confidence score and an absolute geometric refinement vector. Extensive experiments on the NAVSIM benchmark show that DriveVer significantly improves the performance of base planning models. Notably, as an extremely compact model with only 34M parameters, DriveVer introduces minimal computational overhead, achieving competitive results while maintaining real-time inference efficiency.  \nI. INTRODUCTION  \nIn recent years, end-to-end autonomous driving has shown remarkable progress in handling complex urban driving scenarios [1], [2] . To improve robustness against long-tail and rare corner cases, current research primarily relies on training-time scaling, increasing model capacity and training data to enhance planning performance. However, this strategy incurs substantial computational cost while yielding diminishing performance gains as model scale continues to grow.  \nMore importantly, most existing end-to-end planners adopt a one-shot generation paradigm [3], [4], [5], [6], [7], [8],[9], where the planning trajectory is generated once from sensor observations and directly executed by the downstream controller. Consequently, the planner cannot verify or revise its prediction during inference. When encountering complex or adversarial scenarios, suboptimal trajectory predictions cannot be effectively detected or corrected, limiting the safety and reliability of autonomous driving systems.  \nRecently, test-time scaling, also referred to as test-time compute, has emerged as a promising paradigm for improving model performance through additional inference-time computation (e.g., in large language models [10]) . Rather than relying solely on larger models, it enables iterative verification and self-correction during inference. However,  \n1Tsinghua University, 2University of Macau  \n*Equal contribution. ✉ Corresponding author.  \nFig. 1: Conceptual comparison of different planning paradigms. (a) The base planner directly outputs a oneshot trajectory for execution, which can be unsafe. (b) Our proposed DriveVer refines this initial trajectory at test time, providing a safety score and a corrected trajectory.  \ndirectly applying this paradigm to autonomous driving remains challenging due to stringent real-time constraints, as existing inference mechanisms often introduce computational overhead incompatible with the millisecond-level latency requirements of end-to-end planning [3] .  \nTo address this gap and overcome the limitations of pure training-time scaling, we propose DriveVer, a lightweight post-processing framework for trajectory evaluation and refinement. As illustrated in Fig. 1, we compare the conventional planning pipeline with our verification paradigm. Unlike existing end-to-end planners that directly execute a one-shot trajec","cbCaiebqDcntXM7K","https://ap.wps.com/l/cbCaiebqDcntXM7K","pdf",2330871,4,1,"English","en",105,"# Introduction\n## Motivation and limitations of one-shot planning\n## Test-time scaling and verification paradigm\n## Proposed DriveVer framework and contributions","[{\"question\":\"Why do end-to-end autonomous driving models face limitations under training-time scaling?\",\"answer\":\"Training-time scaling raises computational cost and offers diminishing marginal gains, which limits practical improvements in planning performance.\"},{\"question\":\"What problem does DriveVer address compared with one-shot trajectory planners?\",\"answer\":\"DriveVer adds an inference-time safety-critical verification stage so the system can estimate trajectory quality and refine suboptimal trajectories instead of executing a single prediction blindly.\"},{\"question\":\"How does DriveVer remain lightweight while improving safety and reliability?\",\"answer\":\"It uses a lightweight plug-and-play dual-head architecture that predicts a safety confidence score and a geometric refinement vector, reaching competitive results with only 34M parameters and minimal computational overhead for real-time inference.\"}]",1784188320,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"drivever-lightweight-trajectory-evaluator-as-test-time-verifier-for-autonomous-driving","",{"@graph":35,"@context":84},[36,52,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":20},"https://docshare.wps.com/document/drivever-lightweight-trajectory-evaluator-as-test-time-verifier-for-autonomous-driving/83482/",{"url":51,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":23,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"Why do end-to-end autonomous driving models face limitations under training-time scaling?","Question",{"text":74,"@type":75},"Training-time scaling raises computational cost and offers diminishing marginal gains, which limits practical improvements in planning performance.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"What problem does DriveVer address compared with one-shot trajectory planners?",{"text":79,"@type":75},"DriveVer adds an inference-time safety-critical verification stage so the system can estimate trajectory quality and refine suboptimal trajectories instead of executing a single prediction blindly.",{"name":81,"@type":72,"acceptedAnswer":82},"How does DriveVer remain lightweight while improving safety and reliability?",{"text":83,"@type":75},"It uses a lightweight plug-and-play dual-head architecture that predicts a safety confidence score and a geometric refinement vector, reaching competitive results with only 34M parameters and minimal computational overhead for real-time inference.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":28,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":28,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]