[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82534-en":3,"doc-seo-82534-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82534,549758146520,"Patrick","https://ap-avatar.wpscdn.com/avatar/80002397d8c0411e94?_k=1775819394049821470",8,"Research & Report","Flow-Map GRPO: Reinforcement Learning for Few-Step Flow-Map Generators via Anchored Stochastic Composition","Few-step flow-map generators accelerate sampling by learning long-range transport maps between noise and data, yet they are typically deterministic, making reinforcement learning (RL) post-training difficult because it relies on stochastic trajectories and well-defined likelihood ratios. Existing SDE-based stochasticization targets velocity-based samplers and does not transfer directly to long-range flow maps. Flow-Map GRPO introduces Anchored Stochastic Flow Map Composition for path-preserving stochasticization, enabling online RL post-training for deterministic generators without changing their parameterization.","arXiv :2607 .00535v 1 [ cs .LG] 1 Jul 2026  \nFLOW-MAP GRPO: REINFORCEMENT LEARNING FOR FEW-STEP FLOW-MAP GENERATORS VIA ANCHORED STOCHASTIC COMPOSITION  \nZhiqi Li 1 Wen Zhang 1 Bo Zhu 1  \nABSTRACT  \nFew-step flow-map generators, such as consistency models and MeanFlow, accelerate sampling by directly learning long-range transport maps between noise and data. However, these models are typically deterministic, which makes them difficult to optimize with reinforcement learning (RL) post-training methods that require stochastic trajectories and well-defined likelihood ratios. Existing SDEbased stochasticization techniques are designed for velocity-based samplers with infinitesimal or finely discretized transitions, and therefore do not directly apply to long-range flow maps. In this work, we propose Flow-Map GRPO, an online RL post-training framework for deterministic few-step flow-map generators. The key component is Anchored Stochastic Flow Map Composition (ASFMC), a path-preserving stochasticization mechanism that introduces randomness through anchor-based conditional resampling while preserving the original marginal probability path of the deterministic flow map. We derive GRPO objectives for both single-time and two-time flow-map parameterizations. Experiments on few-step FLUX-based text-to-image generators, including MeanFlow and sCM, show that Flow-Map GRPO improves pretrained deterministic flowmap models across reward-based, perceptual, and task-level evaluation metrics.  \nOur results demonstrate that deterministic few-step flow-map generators can be effectively aligned with RL post-training without modifying their original model parameterization or retraining them as native stochastic models.  \n1 INTRODUCTION  \nDiffusion models and continuous-time flow-based generative models have become a dominant paradigm for high-quality image and video generation Rombach et al. (2022); Ho et al. (2022); Dao et al. (2023) . These methods construct a probability path between a simple prior distribution and the data distribution, and learn either a score field or a velocity field that defines a continuous-time generative process. In particular, flow-based approaches such as Flow Matching and Rectified Flow represent generation through a probability-flow ODE, which enables principled sampling by numerically integrating the learned velocity field. Despite their strong theoretical grounding, however, ODE-based sampling typically requires many discretization steps, leading to high computational cost at inference time.  \nThis has motivated recent advances in few-step flow-map-based generative models, such as Consistency Models Song et al. (2023); Geng et al. (2024) and MeanFlow Geng et al. (2025a;b); Li et al.(2026) . Instead of learning only the instantaneous dynamics, these methods directly learn long-range mappings ψt→r that map samples between two time points. By amortizing numerical integration into learned long-range mappings, flow-map-based models can replace iterative ODE solvers with one-step or few-step generation.  \nHowever, deterministic few-step flow maps pose a difficulty for RL post-training, which aims to align a pretrained generator with task-level rewards while preserving its original flow-map parameterization and learned marginal probability path. Recent reinforcement learning (RL) post-training methods for generative models, such as DDPO Black et al. (2024) and Flow-GRPO Liu et al. (2026),  \n1 Georgia Institute of Technology, Atlanta, GA, USA. Correspondence to: Zhiqi Li \u003C [zli3167@gatech.edu](zli3167@gatech.edu)>.  \nformulate sampling as a Markov decision process and optimize the generative policy using task-level rewards. A key requirement of these methods is a well-defined stochastic transition kernel, which is needed both for trajectory-level exploration and for computing likelihood ratios in policy-gradient optimization. Diffusion models naturally provide such stochastic transitions through their denoising process, wh","cbCaiiSUuIS7BcBy","https://ap.wps.com/l/cbCaiiSUuIS7BcBy","pdf",24331674,2,1,31,"English","en",105,"# Abstract\n# Introduction\n## Background: flow-based and diffusion generative models\n## Motivation: RL post-training mismatch for deterministic flow maps\n## Proposed approach: Flow-Map GRPO and ASFMC","[{\"question\":\"Does Flow-Map GRPO require retraining the generator as a native stochastic model or changing its parameterization?\",\"answer\":\"No. Flow-Map GRPO aligns deterministic few-step flow-map generators with RL post-training effectively without modifying their original model parameterization or retraining them as native stochastic models.\"}]",1784181339,78,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"flow-map-grpo-reinforcement-learning-for-few-step-flow-map-generators-via-anchored-stochastic-composition","",{"@graph":36,"@context":77},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/flow-map-grpo-reinforcement-learning-for-few-step-flow-map-generators-via-anchored-stochastic-composition/82534/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"Does Flow-Map GRPO require retraining the generator as a native stochastic model or changing its parameterization?","Question",{"text":75,"@type":76},"No. Flow-Map GRPO aligns deterministic few-step flow-map generators with RL post-training effectively without modifying their original model parameterization or retraining them as native stochastic models.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":98,"slug":130},19,"General","general"]