[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83087-en":3,"doc-seo-83087-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83087,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","Training-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement","Vision-Language-Action (VLA) models enable generalizable robotic manipulation by mapping visual observations and language instructions to smooth, often multimodal action trajectories. Flow matching-based VLA policies succeed but suffer from an expensive iterative denoising process in the action head, limiting real-time deployment. ActionCache introduces a training-free, plug-and-play external cache that reuses past intermediate action chunks via compact multimodal keys to warm-start generation near target actions. Simulation and real-world experiments show strong latency–success improvements, reaching up to 11.75× and 34.43× acceleration for representative flow-based models.","arXiv :2607 .06370v 1 [ cs .RO] 7 Jul 2026  \nTraining-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement  \nRyuji Oi∗ ,† Hikari Otsuka∗ Kosuke Matsushima∗ Yuki Ichikawa  \nMasato Motomura Tatsuya Kaneko Daichi Fujiki  \nInstitute of Science Tokyo  \n†[oi.ryuji@artic.iir.isct.ac.jp](oi.ryuji@artic.iir.isct.ac.jp)  \nAbstract  \nVision-Language-Action (VLA) models have emerged as a promising approach for generalizable robotic manipulations. In particular, flow matching-based VLA models have shown remarkable success due to their capability to generate precise and smooth action sequences and capture multimodal distributions. However, the iterative denoising process in the action head acts as a major computational bottleneck, posing a critical challenge for real-time deployment. To address this challenge, we propose ActionCache, a plug-and-play external cache that opportunistically reuses past intermediate actions to warm-start generations from the vicinity of target actions, thereby drastically reducing the inference latency. Specifically, ActionCache stores the intermediate actions with compact multimodal keys, which enables retrieval from similar past contexts across different episodes or even different tasks.  \nExperimental results in simulation and real-world environments demonstrate that ActionCache maintains high task success rates in a low-latency regime, achieving inference acceleration of up to 11.75 × and 34.43 × for representative flow-based VLA models, π0.5 and GR00T-N1.6, respectively.  \n1 Introduction  \nVision-Language-Action (VLA) models have emerged as a promising foundation for generalist robot policies that map visual observations and language instructions directly to low-level control [7, 12, 15, 16, 20, 24, 35] . For robot manipulation, these policies must convert semantic task understanding into precise, smooth, and often multimodal action trajectories. This requirement has motivated recent VLAs to adopt diffusion-or flow-based action heads, which iteratively produce continuous action chunks rather than autoregressive sequences of discretized action tokens [3–5, 26] . By operating directly in a continuous action space, it avoids action quantization artifacts, captures multimodal action distributions, and produces smooth short-horizon trajectories suitable for closed-loop control. Despite these advantages, iterative action generation introduces a practical bottleneck for real-time robot control. In diffusion-family VLA policies, the action head must repeatedly evaluate a generative model to transform noise into a structured action trajectory. These repeated evaluations can account for a substantial fraction of control-loop latency (e.g., over 65% for DreamVLA [34] and 36% for π0 [5]), especially when the robot replans at high frequency. Reducing the complexity or the number of denoising steps is a natural way to accelerate inference, while affecting the quality and latency trade-off of the policy [31] . Simple tasks may remain solvable with very few refinement steps, whereas harder manipulation tasks often require sufficient refinement to maintain a high success rate. More recently, warm-start methods have exploited temporal continuity by initializing refinement  \nfrom recent actions, predicted actions, or trajectory-level priors [9, 11, 14, 17] . These methods are ∗ Equal contribution.  \nPreprint.  \ntypically framed as improving local temporal smoothness by breaking the Markov Decision Process of VLAs, where the model is conditioned solely on the current observation. As a byproduct, they can also shorten the effective distance traveled during refinement by starting closer to the target action trajectory, rather than forcing the model to synthesize an action chunk from noise. However, this effect has not been systematically studied as a general mechanism for output reuse beyond temporal continuity. Moreover, existing warm-start methods often depend on learned predictors or explicit","cbCaicV0YUUtbLab","https://ap.wps.com/l/cbCaicV0YUUtbLab","pdf",2232327,3,1,14,"English","en",105,"# Introduction\n# Background and Related Work","[{\"question\":\"What problem does ActionCache target in flow-based VLA models?\",\"answer\":\"It targets the iterative denoising in the action head, which becomes a major computational bottleneck and increases inference latency for real-time robot control.\"},{\"question\":\"How does ActionCache reduce inference latency without retraining?\",\"answer\":\"ActionCache caches intermediate action chunks and retrieves them using compact multimodal keys so refinement starts closer to the conditional target flow, reducing the number of required denoising steps in a training-free manner.\"},{\"question\":\"What experimental results does the document report?\",\"answer\":\"Experiments in simulation and real-world settings show ActionCache maintains high task success in low-latency regimes, achieving acceleration up to 11.75× for π0.5 and 34.43× for GR00T-N1.6.\"}]",1784185104,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"training-free-acceleration-for-vision-language-action-models-with-action-caching-and-refinement","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/training-free-acceleration-for-vision-language-action-models-with-action-caching-and-refinement/83087/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does ActionCache target in flow-based VLA models?","Question",{"text":75,"@type":76},"It targets the iterative denoising in the action head, which becomes a major computational bottleneck and increases inference latency for real-time robot control.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does ActionCache reduce inference latency without retraining?",{"text":80,"@type":76},"ActionCache caches intermediate action chunks and retrieves them using compact multimodal keys so refinement starts closer to the conditional target flow, reducing the number of required denoising steps in a training-free manner.",{"name":82,"@type":73,"acceptedAnswer":83},"What experimental results does the document report?",{"text":84,"@type":76},"Experiments in simulation and real-world settings show ActionCache maintains high task success in low-latency regimes, achieving acceleration up to 11.75× for π0.5 and 34.43× for GR00T-N1.6.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]