[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82452-en":3,"doc-seo-82452-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82452,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","OpenLongTail: Generative Scaling of Long-Tail Driving Data","OpenLongTail addresses the bottleneck in scaling robust autonomous driving policies caused by scarce edge cases in curated datasets. Real-world long-tail events are captured continuously but remain underused because heterogeneous sources produce long-tail videos with a modality gap: missing multi-view coverage, absent calibrated multi-camera poses, and incompatible annotations. OpenLongTail provides an open-source generative data engine that synthesizes missing, pose-aligned multi-view assets using pose-informed extrapolative view synthesis and Plücker ray geometry, improving closed-loop robustness and validating visual fidelity, cross-view consistency, and ego-trajectory recovery.","arXiv :2607 .09655v1 [ cs .CV] 10 Jul 2026  \nOpenLongTail: Generative Scaling of Long-Tail Driving Data  \nLulin Liu∗1 , Nuo Chen∗1 , Yan Wang2 , Bangya Liu3 , Wenyan Cong4 , Hezhen Hu4 , Boris Ivanovic2 , Hao Wang 1 , Ziyao Zeng5 , Xinyu Gong6 , Yang Zhou 1 , Zixiang Xiong 1 , Dilin Wang7 , Zhangyang Wang4 , Weisong Shi8 , Ruohan Zhang9 , Marco Pavone2 ,9 , Zhiwen Fan 1 ,†  \n1 Texas A&M University, 2 NVIDIA, 3 UW–Madison, 4 UT Austin, 5Yale University, 6 Adobe, 7 Meta, 8 University of Delaware, 9 Stanford University  \n∗ Equal contribution, †Corresponding author  \nScaling robust driving policies is fundamentally bottlenecked by the scarcity of edge cases in curated datasets. While the real world continuously captures these critical events, such long-tail events remain underutilized when collected from heterogeneous sources. Specifically, diverse but valuable in-the-wild long-tail videos lack the full view coverage required for training policy models, often missing multi-view poses or originating solely from monocular dash cameras. This modality gap prevents these ubiquitous observations from being converted into scalable training data for long-tail generalization. We introduce OpenLongTail, an open-source generative data engine for scaling autonomous driving policies under long-tail events. To transform heterogeneous data sources into view-aligned and temporally coherent multi-view assets that are useful for policy learning, we develop a pose-informed extrapolative view synthesis pipeline that generates the missing views. We further enhance cross-view consistency and the temporal alignment for the newly generated views by injecting Plücker ray geometry into the scalable generation engine. By synthesizing heterogeneous long-tail data, we observe a significant improvement in closed-loop driving robustness in handling long-tail events. By measuring the extrapolative view synthesis and pose metrics, we validate the effectiveness of OpenLongTail in visual fidelity, cross-view consistency, and ego-trajectory recovery.  \nDate: July 13, 2026  \nKeywords: Autonomous Driving, Video Diffusion Model, VLA Model  \nProject Website: [https://openlongtail.github.io/](https://openlongtail.github.io/)  \n1 Introduction  \nEmpowered by large-scale driving datasets (Sun et al. , 2020 ; Ettinger et al. , 2021 ; Caesar et al. , 2020 ; Xu et al. , 2025), recent Vision-Language-Action (VLA) driving policies have made substantial progress on common scenarios by learning end-to-end mappings from scene understanding to vehicle control under well-structured training conditions (Wang et al. , 2025c ; Zhou et al. , 2025 , 2026 ; Xu et al. , 2024a ; Yuan et al. , 2025 ; Jiang et al. , 2025b ; Hwang et al. , 2024 ; Jiang et al. , 2024 ; Feng et al. , 2025 ; Jiang et al. , 2025a) . However, real-world autonomy also requires these policies to remain reliable in long-tail situations, where dynamic obstacles and atypical environmental conditions might lead to safety-critical failures. Yet several challenges arise for the training data. First, events like animals on the road and work zones are scarce in curated datasets and are usually more expensive to capture at scale with calibrated multi-camera rigs. Sometimes though such events are recorded, the resulting data is fragmented across heterogeneous sources, spanning from recently released large-scale repositories such as the NVIDIA PhysicalAI Autonomous Vehicles (PAV) (NVIDIA Corporation, 2025) and Waymo E2E dataset (Xu et al. , 2025) to widely available monocular dash-camera videos. These observations often follow incompatible annotation formats and provide missing or unreliable metric camera poses. This modality gap prevents common real-world long-tail videos from being directly converted into synchronized multi-view training assets, limiting the continued scaling of learned driving policies toward robust long-tail generalization.  \nBridging this gap requires a conversion engine that turns unposed or monocular driv","cbCaih4ynfoav5eO","https://ap.wps.com/l/cbCaih4ynfoav5eO","pdf",14995221,3,1,25,"English","en",105,"# Introduction\n## Problem: Long-tail data scarcity and modality gap\n## Approach: Pose-informed extrapolative view synthesis\n## Contribution and validation strategy","[{\"question\":\"What problem does OpenLongTail target in long-tail autonomous driving training data?\",\"answer\":\"It targets the difficulty of scaling robust driving policies because curated datasets lack rare edge cases and because real long-tail videos cannot be converted into synchronized multi-view assets due to modality gaps and missing pose/coverage.\"},{\"question\":\"How does OpenLongTail generate missing views for heterogeneous long-tail videos?\",\"answer\":\"It uses a pose-informed extrapolative view synthesis pipeline to create view-aligned, temporally coherent multi-view assets that match a target camera rig, generating views beyond the observed front camera.\"},{\"question\":\"How does OpenLongTail improve consistency and temporal alignment in the synthesized views?\",\"answer\":\"It injects Plücker ray geometry into the scalable generation engine, enhancing cross-view consistency and aligning the newly generated views over time.\"}]",1784180476,63,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"openlongtail-generative-scaling-of-long-tail-driving-data","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/openlongtail-generative-scaling-of-long-tail-driving-data/82452/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does OpenLongTail target in long-tail autonomous driving training data?","Question",{"text":75,"@type":76},"It targets the difficulty of scaling robust driving policies because curated datasets lack rare edge cases and because real long-tail videos cannot be converted into synchronized multi-view assets due to modality gaps and missing pose/coverage.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does OpenLongTail generate missing views for heterogeneous long-tail videos?",{"text":80,"@type":76},"It uses a pose-informed extrapolative view synthesis pipeline to create view-aligned, temporally coherent multi-view assets that match a target camera rig, generating views beyond the observed front camera.",{"name":82,"@type":73,"acceptedAnswer":83},"How does OpenLongTail improve consistency and temporal alignment in the synthesized views?",{"text":84,"@type":76},"It injects Plücker ray geometry into the scalable generation engine, enhancing cross-view consistency and aligning the newly generated views over time.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]