[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85894-en":3,"doc-seo-85894-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85894,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","VINE: Taming Generative Control Policies for Reinforcement Learning","Flow-matching policies provide expressive, multimodal action generation for robot learning by iteratively denoising from noise. Scaling them with value-gradient reinforcement learning (RL) often causes severe training instability, previously attributed to the iterative generation process. This work shows the instability instead originates from the vanilla behavior-cloning sampling strategy, which becomes brittle under value-gradient RL. VINE introduces an RL-oriented sampling method that reconstructs an interpolation state at every denoising step, enabling stable end-to-end value-gradient propagation while preserving flow-matching denoising and expressiveness, achieving robust improvements on OGBench offline RL and real robotic manipulation.","arXiv :2607 . 10369v1 [ cs .RO] 11 Jul 2026  \nVINE: Taming Generative Control Policies for Reinforcement Learning  \nRushuai Yang1,2 Zhuo Han1 Houlin Li1 Hecheng Wang1 Zhichao Wu 1 Rui Zhang 1 Zhaowei Zhang3 Zihong Chen 1 Xiaohan Yan 1 Chiming Liu 1,† Yi Chen2 Wei Shan 1,† Maoqing Yao 1,†  \n1AgiBot 2The Hong Kong University of Science and Technology 3Peking University  \n†Corresponding Author  \nAbstract: Flow-matching policies have emerged as an effective policy parameterization for robot learning. They iteratively generate actions from noise, enabling highly expressive modeling of complex and multimodal action distributions. However, prior works observed that scaling these policies with valuegradient reinforcement learning (RL) often leads to training instability. Existing methods attribute this instability to iterative generation and therefore avoid end-to-end value-gradient optimization by sacrificing iterative generation, high expressiveness, or value-gradient optimization. Contrary to prior belief, we show the instability does not stem from iterative generation itself, but from the vanilla sampling strategy originally designed for behavior cloning, which becomes brittle under value-gradient RL. Motivated by this insight, we propose VINE, an RL-oriented sampling method that enables stable end-to-end value-gradient optimization for flow-matching policies. Instead of following a single flow trajectory, VINE reconstructs a new interpolation state at every denoising step, creating astable differentiable path for value-gradient propagation while remaining compatible with the original flow-matching denoising process. As a result, VINE preserves the expressiveness and iterative generation of flow-matching without sacrificing end-to-end value-gradient optimization. Despite performing end-toend backpropagation through all ten denoising steps, VINE achieves stable policy improvement and consistently outperforms state-of-the-art RL methods on the OGBench offline RL benchmark and real-world robotic manipulation task. Videos are available on our website: [https://agibottech.github.io/vine](https://agibottech.github.io/vine).  \n\"!  \n!!  \nBehavior Flow Policy  \n!!  \n\n|  |  |\n| --- | --- |\n\n~~ ~~ Ideal ~~ ~~ Sampling path  Gradient chain  Q-gradient direction  Drifted BPTT path  Correct path  \nFigure 1: Left: A flow policy trained by behavior cloning fits the multi-modal data distribution. Middle: Directly backpropagating the critic gradient through the denoising steps (value BPTT) destabilizes the trajectory. Right: VINE produces a stable denoising trajectory that supports valuegradient BPTT toward a 1.  \n1 Introduction  \nIn robotic learning, generative control policies such as diffusion and flow-matching models have achieved remarkable success for behavior cloning in a wide range of manipulation and control tasks [1, 2, 3, 4] . In general, they start from randomly sampled noise, and iteratively refine the noise using a learned velocity or score field until an executable action is produced. By combining structured noise injection with iterative refinement, these policies can represent complex and multimodal action distributions with substantially greater expressive power than conventional policy parameterizations [2] . Because human demonstrations are not always optimal and do not directly align with reward maximization, prior work has begun to further improve these generative control policies with reinforcement learning [5, 6, 7, 8, 9, 10] . In this paradigm, the policy is trained to maximize expected returns, using a learned value function to estimate future returns as the optimization signal. Yet prior work has found that such approaches suffer from severe training instability when value-gradient optimization is applied to highly expressive generative policies [11, 12, 13, 14, 15] .  \nTo mitigate this problem, existing work generally attributes this instability to the iterative generation process, and therefore avoids propagating value ","cbCailYisgS5LYOn","https://ap.wps.com/l/cbCailYisgS5LYOn","pdf",4886097,6,1,20,"English","en",105,"# Introduction\n## Background: generative control policies and instability\n## Prior mitigation strategies and their trade-offs\n## Key insight: sampling strategy as the instability bottleneck","[{\"question\":\"What problem does VINE address in reinforcement learning for flow-matching control policies?\",\"answer\":\"Scaling flow-matching policies with value-gradient RL leads to severe training instability, which prior work tried to avoid by limiting end-to-end value-gradient propagation.\"},{\"question\":\"Why is the instability not caused by iterative denoising itself?\",\"answer\":\"The document argues that the vanilla sampling strategy, originally designed for behavior cloning, becomes brittle under value-gradient RL, rather than the iterative generation process being the root cause.\"},{\"question\":\"How does VINE enable stable end-to-end value-gradient optimization?\",\"answer\":\"VINE reconstructs a new interpolation state at every denoising step, forming a stable differentiable path for value-gradient propagation while remaining compatible with the original flow-matching denoising process.\"}]",1784207006,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"vine-taming-generative-control-policies-for-reinforcement-learning","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/vine-taming-generative-control-policies-for-reinforcement-learning/85894/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does VINE address in reinforcement learning for flow-matching control policies?","Question",{"text":76,"@type":77},"Scaling flow-matching policies with value-gradient RL leads to severe training instability, which prior work tried to avoid by limiting end-to-end value-gradient propagation.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"Why is the instability not caused by iterative denoising itself?",{"text":81,"@type":77},"The document argues that the vanilla sampling strategy, originally designed for behavior cloning, becomes brittle under value-gradient RL, rather than the iterative generation process being the root cause.",{"name":83,"@type":74,"acceptedAnswer":84},"How does VINE enable stable end-to-end value-gradient optimization?",{"text":85,"@type":77},"VINE reconstructs a new interpolation state at every denoising step, forming a stable differentiable path for value-gradient propagation while remaining compatible with the original flow-matching denoising process.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":29,"slug":113},"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":22,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":22,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":107,"slug":136},19,"General","general"]