[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82162-en":3,"doc-seo-82162-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82162,687197207057,"Sage","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Can the Cloud Drive? Infrastructure Feasibility of Offloading Autonomous Driving Across 5G and 6G","Frontier autonomous-driving models—especially vision-language-action (VLA) models with forward passes near 60 TFLOPs—outgrow cost-effective onboard deployment because peak hardware remains idle most of the day. Cloud inference can pool GPUs across active vehicles, yet requires capacity-limited uplink upload, queue-free GPU access, and closed-loop decision return within latency budgets. A coupled analytical framework links communication limits, roofline GPU service, stochastic latency, and utilization-aware costs for NYC.","Highlights  \nCan the Cloud Drive? Infrastructure Feasibility of Offloading Autonomous Driving Across 5G and 6G  \nPouya Parsa, Kawon Han, Seongjin Choi  \n• Framework links AV model complexity, 5G/6G limits, and cloud GPU economics  \n• Three nested NYC regimes: communication binds, then VLA compute, then cost  \n• Feature-level offloading (S2) is where the VLA cloud cost crossover concentrates  \n• Latency sets which model is admissible each year; cost sets if it is economical  \narXiv :2607 .09045v1 [ ee ss . SY] 10 Jul 2026  \nCan the Cloud Drive? Infrastructure Feasibility of Offloading Autonomous Driving Across 5G and 6G  \nPouya Parsaa , Kawon Hanb and Seongjin Choia,∗  \na Department of Civil, Environmental, and Geo-Engineering, University of Minnesota, Twin Cities, Minneapolis, 55455, Minnesota, United States b Department of Electrical Engineering, Ulsan National Institute of Science and Technology (UNIST), Ulsan, 44919, South Korea  \nARTICLE INFO  \nKeywords:  \nVehicular edge computing Task offloading Autonomous driving 5G/6G communications Infrastructure feasibility  \nAB STRACT  \nFrontier autonomous-drivi∼ng models—especially vision-language-action (VLA) models, whose  \nforward pass approaches 60 TFLOPs—are outgrowing economical onboard deployment, since peak hardware sits idle most of the day. Cloud inference can instead share GPUs across active vehicles, but the vehicle must upload through a capacity-limited uplink, reach a GPU without queueing, and return a decision within the closed-loop budget. This paper asks: can the cloud drive? We answer with an analytical framework coupling communication limits, a roofline GPU service model, stochastic latency, and utilization-aware cost across three model classes, three offloading strategies, and three communication generations, applied to New York City. Separating a reactive 100 ms budget from a 300 ms deliberative tier (presuming an onboard reactive fallback), we find three nested binding regimes. Communication binds first in dense cells: 5G fails early, 5G-Advanced is the practical threshold for feature-level offloading, and 6G adds headroom. Compute binds next under the reactive budget: near-term VLA is latencyinfeasible∼ regardless of bandwidth, because autoregressive FP16 decode is memory-bandwidthbound ( 114 ms on 20∼25 hardware) . Its floor clears 100 ms around 2027; 6G then admits  \nfeature-level VLA by 2028, 5G-Advanced only at light loading and not the dense corridor, and the deliberative tier from 2026 . Cost binds last: once admissible, utilization-pooled cloud GPUs undercut onboard hardware for VLA, whose baseline (up to $8,500 per vehicle-year) is expensive and idle; feature-level offloading (S2) is where the VLA cost crossover concentrates. Latency decides which model is admissible in which year; cost decides whether it is economical.  \n1. Introduction  \nAutonomous driving (AD) is moving from modular perception-prediction-planning pipelines toward end-to-end models that map sensor inputs directly to driving actions. This transition has widened the computational range of AD models substantially. Compact end-to-end (E2E) architectures such as UniAD [1] and VAD [2] operate around a few TFLOPs (tera floating-point operations) per forward pass; Vision Language Model (VLM)-based models such as DriveLM [3] add language-conditioned reasoning; and recent Vision-Language-Action (VLA) models such as EMMA [4] and Alpamayo [5] push perception, reasoning, and action into a single large model. While current automotive processors can marginally accommodate E2E and VLM workloads, running larger models such as VLAs remains a practical deployment challenge.  \nThe straightforward response is to put more hardware in every vehicle. For personal vehicles, however, this approach is economically inefficient since each vehicle has to be equipped with hardware sized for its peak inference load. The peak arises only during rare driving moments when the full autonomous-driving pipeline must run ","cbCaiumKL87FC6XK","https://ap.wps.com/l/cbCaiumKL87FC6XK","pdf",1026952,1,26,"English","en",105,"# Introduction\n## Motivation: From onboard autonomy to cloud inference\n## Central question: Can the cloud drive?\n## Joint constraints: model split, communication, and cloud GPU service","[{\"question\":\"Why are large autonomous-driving models difficult to deploy onboard today?\",\"answer\":\"VLA and related end-to-end models have very high forward-pass compute (tens of TFLOPs), while vehicle hardware is mostly idle about 95% of the day, making peak-sized onboard deployment uneconomical.\"},{\"question\":\"What system constraints must be satisfied for cloud offloading to work?\",\"answer\":\"Cloud driving depends on three coupled components: the driving-model split (what runs onboard vs. in the cloud), vehicle-to-network communication limits for continuous uploads, and cloud GPU service that must meet the closed-loop deadline without queueing.\"},{\"question\":\"How do latency and cost jointly determine which models are feasible over time?\",\"answer\":\"Latency determines whether a model class is admissible in a given year under reactive vs deliberative budgets, while cost determines whether pooled cloud GPUs undercut onboard hardware; the feature-level offloading strategy concentrates the cost crossover for VLA.\"}]",1784178531,66,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"can-the-cloud-drive-infrastructure-feasibility-of-offloading-autonomous-driving-across-5g-and-6g","",{"@graph":35,"@context":84},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/can-the-cloud-drive-infrastructure-feasibility-of-offloading-autonomous-driving-across-5g-and-6g/82162/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"Why are large autonomous-driving models difficult to deploy onboard today?","Question",{"text":74,"@type":75},"VLA and related end-to-end models have very high forward-pass compute (tens of TFLOPs), while vehicle hardware is mostly idle about 95% of the day, making peak-sized onboard deployment uneconomical.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"What system constraints must be satisfied for cloud offloading to work?",{"text":79,"@type":75},"Cloud driving depends on three coupled components: the driving-model split (what runs onboard vs. in the cloud), vehicle-to-network communication limits for continuous uploads, and cloud GPU service that must meet the closed-loop deadline without queueing.",{"name":81,"@type":72,"acceptedAnswer":82},"How do latency and cost jointly determine which models are feasible over time?",{"text":83,"@type":75},"Latency determines whether a model class is admissible in a given year under reactive vs deliberative budgets, while cost determines whether pooled cloud GPUs undercut onboard hardware; the feature-level offloading strategy concentrates the cost crossover for VLA.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":105,"slug":137},19,"General","general"]