[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86423-en":3,"doc-seo-86423-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86423,4398048949847,"Eliana","https://ap-avatar.wpscdn.com/avatar/400002536579ef2da7f?_k=1778318612642679267",8,"Research & Report","FAST Framework for Aligned Sampling and Training in Parallel Reinforcement Learning for Autonomous Driving","Deep reinforcement learning is essential for closed-loop autonomous driving but is hindered by severe sampling-efficiency bottlenecks. Conventional parallel sampling alleviates throughput limitations yet suffers from the straggler effect: the earliest termination forces synchronized batch re-initialization, wasting long-horizon samples and incurring costly reset latency. FAST introduces Dynamic Parallel Sampling Alignment to virtual-continue terminated episodes and decouple sampling from individual terminations. Scaled Mask Padding Optimization preserves theoretical unbiasedness by removing bias from auxiliary padding via validity masking and adaptive loss normalization. Experiments show at least 1.78× wall-clock speedup without losing statistical correctness.","FAST: A Framework for Aligned Sampling and Training in Parallel Reinforcement Learning for Autonomous Driving  \nBonan Wang∗ , Letian Tao∗ , Bin Shuai, Jiaxin Gao, Wenxin Zhao, Wei Xiong, Kehua Sheng, Bo Zhang, Yang Guan†, and Shengbo Eben Li† Senior Member, IEEE  \narXiv :2606 .2 1587v2 [ cs .LG] 13 Jul 2026  \nAbstract—Deep reinforcement learning is pivotal for closedloop autonomous driving yet remains constrained by severe bottlenecks in sampling efficiency. Standard parallel sampling mitigates this but suffers from the straggler effect, where the premature termination of a single environment necessitates a synchronized batch re-initialization, leading to suboptimal sample utilization and prohibitive re-initialization latency. To address this, we propose FAST, a synchronous parallel framework tailored for closed loop simulation. Specifically, FAST employs Dynamic Parallel Sampling Alignment (DPSA) to maintain vectorization synchronization by extending terminated episodes via virtual continuation, thereby decoupling the sampling loop from individual terminations. By dynamically triggering global truncation based on the termination rate of parallel clips, FAST effectively eliminates the bottleneck of premature resets without sacrificing data diversity. Furthermore, to strictly preserve theoretical consistency, we incorporate a Scaled Mask Padding Optimization (SMPO) that leverages validity masking and adaptive loss normalization to nullify the bias from auxiliary padding data. Empirical evaluations demonstrate that FAST achieves at least a 1.78× wall-clock speedup over the single-clip baseline while preserving statistical unbiasedness.  \nIndex Terms—Autonomous Driving, Deep Reinforcement Learning, Parallel Sampling, Sampling Efficiency  \nI. INTRODUCTION  \nDEEP Reinforcement Learning (DRL) has emerged as  \na fundamental paradigm in autonomous driving (AD), enabling robust decision-making within high-dimensional and non-stationary environments [1]–[4] . Integrating highdimensional sensory processing with decision-making allows DRL agents to master critical competencies such as path planning, trajectory optimization, and collision avoidance through continuous trial and error. Achieving robust policy performance thus requires training on massive interaction data, often demanding millions or billions of simulation steps to cover long-tail corner cases and ensure safety generalization [5]–[7] . Driven by this immense data demand, the continuous in  \ngestion of high-quality interaction data represents a critical * Authors contributed equally; † Corresponding author.  \nBonan Wang, Letian Tao, Bin Shuai, Jiaxin Gao, Wenxin Zhao and Yang Guan are with the School of Vehicle and Mobility, Tsinghua University, Beijing 100084, China.  \nWei Xiong, Kehua Sheng and Bo Zhang are with Didi Voyager Labs, DiDi Autonomous Driving.  \nShengbo Eben Li are with the School of Vehicle and Mobility and the College of AI, Tsinghua University, Beijing 100084, China.  \nE-mail: [lishbo@tsinghua.edu.cn](lishbo@tsinghua.edu.cn), [yguan@tsinghua.edu.cn](yguan@tsinghua.edu.cn)  \nThis work is supported by the National Natural Science Foundation of China with 92582205 and the Key Program of the Beijing Municipal Natural Science Foundation with L257002 .  \nThis work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible.  \nFig. 1. Illustration of the inefficiency caused by the “reset-on-anytermination” protocol. The global reset is triggered by the earliest terminating clip (clip 1), forcing the premature truncation of surviving clips. This mechanism discards valuable long-horizon data and increases the frequency of computationally expensive re-initializations.  \ncomponent within the pipeline of DRL [8], [9] . In contrast to open-loop paradigms that leverage static, pre-recorded datasets for model optimization, closed-loop simulation necessitates the dynamic generation ","cbCaiuLbHlBWxViD","https://ap.wps.com/l/cbCaiuLbHlBWxViD","pdf",9619438,7,1,11,"English","en",105,"# Introduction\n## Deep reinforcement learning in autonomous driving\n## Sampling bottlenecks and the straggler effect\n## On-policy vs off-policy under synchronization constraints","[{\"question\":\"What problem does FAST address in parallel reinforcement learning for autonomous driving?\",\"answer\":\"FAST targets the sampling-efficiency bottleneck caused by stragglers in synchronized parallel sampling, where the earliest episode termination forces premature truncation of other episodes and costly re-initializations.\"},{\"question\":\"How does Dynamic Parallel Sampling Alignment (DPSA) work in FAST?\",\"answer\":\"DPSA maintains vectorization synchronization by extending terminated episodes through virtual continuation, effectively decoupling the sampling loop from individual termination events.\"},{\"question\":\"How does FAST ensure theoretical consistency when using auxiliary padding?\",\"answer\":\"FAST incorporates Scaled Mask Padding Optimization (SMPO) using validity masking and adaptive loss normalization to nullify bias introduced by auxiliary padding data.\"}]",1784211663,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"fast-framework-for-aligned-sampling-and-training-in-parallel-reinforcement-learning-for-autonomous-driving","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/fast-framework-for-aligned-sampling-and-training-in-parallel-reinforcement-learning-for-autonomous-driving/86423/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does FAST address in parallel reinforcement learning for autonomous driving?","Question",{"text":76,"@type":77},"FAST targets the sampling-efficiency bottleneck caused by stragglers in synchronized parallel sampling, where the earliest episode termination forces premature truncation of other episodes and costly re-initializations.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does Dynamic Parallel Sampling Alignment (DPSA) work in FAST?",{"text":81,"@type":77},"DPSA maintains vectorization synchronization by extending terminated episodes through virtual continuation, effectively decoupling the sampling loop from individual termination events.",{"name":83,"@type":74,"acceptedAnswer":84},"How does FAST ensure theoretical consistency when using auxiliary padding?",{"text":85,"@type":77},"FAST incorporates Scaled Mask Padding Optimization (SMPO) using validity masking and adaptive loss normalization to nullify bias introduced by auxiliary padding data.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]