[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83068-en":3,"doc-seo-83068-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83068,1374391975076,"Riley","https://ap-avatar.wpscdn.com/avatar/14000253ca4ec9f6853?x-image-process=image/resize,m_fixed,w_180,h_180&k=1783305029341752051",8,"Research & Report","Optimal Transport Q-Learning for Flow Policy Steering and Acceleration","Optimal Transport Q-Learning (OTQL) addresses the challenge of improving and accelerating flow-based robot policies when high-quality demonstrations and fast inference are limited. The method fine-tunes flow policies using robot experience through RL post-training, employing advantage-weighted conditional optimal transport flow matching. OTQL avoids computationally expensive distillation in both simulation and real-world tasks, works within a 50–60 episode interaction budget, and reduces inference steps per action generation by about 70%.","arXiv :2607 .06262v 1 [ cs .RO] 7 Jul 2026  \nOptimal Transport Q-Learning for Flow Policy Steering and Acceleration  \nAndreas Sochopoulos 1 , 2 , Esmeralda S. Whitammer 1 , Nikolaos Tsagkas 1 , Joo Moura 1 ,  \nMichael Gienger2 , Sethu Vijayakumar 1  \n1University of Edinburgh 2Honda Research Institute Europe  \n[ansocho.github.io/otql-flow](ansocho.github.io/otql-flow)  \nAbstract: Diffusion and flow policies have recently demonstrated remarkable performance in robotic applications by accurately capturing multimodal robot trajectory distributions, especially in the context of vision language action (VLA) models. However, high quality policy performance also requires fast inference and high quality demonstrations, which are often hard to get. Lack of these leads to suboptimal policy behaviors and failure under distribution shifts. In this work we address the problem of fine-tuning and accelerating suboptimal flow-based policies using the robot’s experience through RL post-training. We introduce Optimal Transport Q-Learning (OTQL), a new methodology for finetuning flow policies using advantage weighted conditional optimal transport flow matching. OTQL can finetune and accelerate flows with an interaction budget of 50-60 episodes while avoiding computationally expensive distillation in simulation and real-world robot tasks. Our results show that OTQL post-trains flow policies using the robot’s own experience, increasing average success percentage of single-task policies from 36% to 86% and of a pre-trained VLA from 38% to 76% while reducing the number of inference steps per action generation by 70% .  \nKeywords: Flow Matching, Reinforcement Learning, Optimal Transport  \n1 Introduction  \nVision language action models (VLAs) and, more recently, world action models (WAMs) have brought the field of robotics one step closer to generalist policies. These systems use as backbone a large, pre-trained vision language model (VLM) [1, 2] or world model (WM), often in the form of a video model [3, 4], that provides features to an action head. This module is most commonly a flow or diffusion model [4, 3, 5, 6, 7, 8] . Flows trained with flow matching (FM) [9, 10] and diffusion models [11, 12] have proven highly effective for action generation, as they can capture complex, multimodal action distributions [13, 14], such as those demonstrated by humans during teleoperation or kinesthetic teaching.  \nDespite undergoing large-scale pretraining, VLAs and WAMs often struggle with zero-shot generalization to tasks unseen during training. Moreover, single and multi-task flow and diffusion policies frequently fail to handle out-of-distribution environment states, leading to task failures. The standard remedy for these issues is collecting additional demonstrations for the target tasks or states, and performing supervised fine-tuning (SFT) on the newly acquired dataset. However, collecting data through teleoperation is tedious and requires significant human effort, time, and expertise.  \nReinforcement learning (RL) provides the tools for policies to adapt using the robot’s own experience, eliminating the need for human demonstrations [15, 16, 17, 18, 19] . A large body of work has attempted to train flow and diffusion policies in offline RL settings using a static set of rollouts [20, 21, 22, 23, 24, 25] . Although successful in simple to moderate robotic tasks, most of these methods suffer from slow inference because they require integrating an SDE or ODE over multiple steps [23] . Furthermore, since critic training typically involves generating actions from the learned policy, this slow inference propagates to every training step, resulting in slow training. While there  \nhave been attempts to train single-step policies [25, 21], they rely on computationally expensive distillation. In the context of behavior cloning, there has been extensive work on diffusion/flow acceleration [26, 27, 28, 29, 30, 31], but methods that address both acceleration and RL po","cbCaifhD87600c3v","https://ap.wps.com/l/cbCaifhD87600c3v","pdf",4432601,2,1,18,"English","en",105,"# Introduction\n# Background","[{\"question\":\"What problem does OTQL target in flow-based robotic policies?\",\"answer\":\"OTQL targets suboptimal behavior caused by missing high-quality demonstrations and the need for fast inference, especially under distribution shifts.\"},{\"question\":\"How does OTQL fine-tune flow policies?\",\"answer\":\"OTQL uses robot experience for RL post-training with advantage-weighted conditional optimal transport flow matching to steer flows toward high-value action samples.\"},{\"question\":\"What efficiency improvements does OTQL provide?\",\"answer\":\"OTQL reduces neural function evaluations per action generation by about 70% and achieves fine-tuned performance with only 2–3 integration steps for inference within a 50–60 episode interaction budget.\"}]",1784184993,45,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"optimal-transport-q-learning-for-flow-policy-steering-and-acceleration","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/optimal-transport-q-learning-for-flow-policy-steering-and-acceleration/83068/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does OTQL target in flow-based robotic policies?","Question",{"text":75,"@type":76},"OTQL targets suboptimal behavior caused by missing high-quality demonstrations and the need for fast inference, especially under distribution shifts.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does OTQL fine-tune flow policies?",{"text":80,"@type":76},"OTQL uses robot experience for RL post-training with advantage-weighted conditional optimal transport flow matching to steer flows toward high-value action samples.",{"name":82,"@type":73,"acceptedAnswer":83},"What efficiency improvements does OTQL provide?",{"text":84,"@type":76},"OTQL reduces neural function evaluations per action generation by about 70% and achieves fine-tuned performance with only 2–3 integration steps for inference within a 50–60 episode interaction budget.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]