[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82043-en":3,"doc-seo-82043-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82043,7971461740909,"Levi","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Reward Transport: Property Control in Flow Matching via Noise-Space Alignment","Flow matching relies on a coupling that pairs noise vectors with data points, usually treated as a computational detail. This work reframes coupling as an alignment interface that embeds property-controlled structure into the learned flow field. Reward Transport aligns a scalar noise-space coordinate with molecular rewards by using optimal transport during training, then steers generation at inference by varying that coordinate—without an oracle, reward model, gradients, or extra computation. Coupling-preserving alignment yields a Cross-Entropy Method-style truncated distribution and supports monotone logP and QED control on ZINC-250K and GuacaMol.","arXiv :2607 .08781v1 [ cs .LG] 13 Jun 2026  \nReward Transport: Property Control in Flow Matching via Noise-Space Alignment  \nKehan Guo 1 Yili Shen 1 Yujun Zhou 1 Yue Huang 1 Chujie Gao 1 Shiyi Du2 Xiangliang Zhang 1  \n1University of Notre Dame 2 Carnegie Mellon University  \n[kguo2@nd.edu](kguo2@nd.edu) [xzhang33@nd.edu](xzhang33@nd.edu)  \nAbstract  \nThe coupling in flow matching—the rule pairing noise vectors with data points—is typically treated as a computational choice. We show that this coupling can instead serve as an alignment interface: by matching noise and data according to a target molecular property, it embeds controllable structure directly into the learned flow field. Building on this view, we introduce Reward Transport, which uses optimal transport coupling at training time to align a scalar noise-space coordinate with molecular rewards; at inference, varying this coordinate steers the generated distribution without requiring an oracle, reward model, gradient guidance, or additional computation. In the coupling-preserving limit, thresholding this coordinate recovers the Cross-Entropy Method’s truncated reward distribution, providing a principled, continuously adjustable distribution-level control knob.  \nEmpirically, on ZINC-250K and GuacaMol, sweeping the scalar induces monotone control of logP and consistent QED control over its operating range; most tellingly, the same knob produces opposite structural responses for different targets, growing molecules for logP but shrinking them for QED, which rules out a generic size bias. The interface is complementary to classifier-free guidance and conditional flow matching, while a negative result under ϵ-prediction diffusion clarifies where coupling-level alignment is structurally absent. Our code is available at:  \n[https://github.com/KehanGuo2/reward-transport](https://github.com/KehanGuo2/reward-transport).  \n1 Introduction  \nA flow matching model depends on both the vector-field training objective and the coupling: the rule that determines which noise vector is paired with which data point during training. While prior work has carefully studied training objectives [Karras et al., 2022, Esser et al., 2024, Ma et al., 2024], the coupling is often treated as a background implementation choice: independent pairing for simplicity [Lipman et al., 2022], or minibatch optimal transport to straighten trajectories and reduce path crossings [Tong et al., 2023, Pooladian et al., 2023] . In every case, the coupling is chosen to make training easier, not to shape what the model ultimately learns.  \nWe take a different view. The coupling is not a training detail but an alignment interface: it determines which noise samples are associated with which data samples, and therefore whether meaningful data properties become organized in noise space. In most controllable generative models [Ho and Salimans, 2022, Dhariwal and Nichol, 2021], steerability is introduced through explicit inputs, such as class labels, property embeddings, or classifier-free guidance. In flow matching, we show that steerability can also be introduced through the coupling itself. By choosing this assignment according to a target property, the learned flow field inherits a property-aligned organization before any inference-time guidance is applied [Peebles and Xie, 2023, Zeng et al., 2025] .  \nPreprint.  \nMolecules are a natural testbed. They have scalar targets that matter—logP [Wildman and Crippen, 1999] and QED [Bickerton et al., 2012]—discrete variable-length samples that stress naive couplings, and properties cheap to verify without an oracle. We study coupling as an alignment interface in the molecular regime and ask whether choosing it alone is enough to buy controllable generation, using property labels only at training time to build the coupling.  \nWe instantiate the interface as Reward Transport (Figure 1): a property-aligned monotone coupling that sorts noise vectors by a scalar coordinate s and molecul","cbCaibjZs84GwYaT","https://ap.wps.com/l/cbCaibjZs84GwYaT","pdf",2425029,2,1,29,"English","en",105,"# Introduction\n## Coupling as an Alignment Interface\n## Reward Transport (Noise-Space Alignment)\n## Empirical Results and Steering Behavior","[{\"question\":\"What empirical evidence supports the method’s controllability and limits?\",\"answer\":\"On ZINC-250K and GuacaMol, sweeping the scalar induces monotone control of logP and consistent QED control across its operating range. The same knob produces opposite structural responses for different targets, arguing against a generic size bias; a negative result under ε-prediction diffusion clarifies where coupling-level alignment is not structurally present.\"}]",1784177767,73,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"reward-transport-property-control-in-flow-matching-via-noise-space-alignment","",{"@graph":36,"@context":77},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/reward-transport-property-control-in-flow-matching-via-noise-space-alignment/82043/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-20","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"What empirical evidence supports the method’s controllability and limits?","Question",{"text":75,"@type":76},"On ZINC-250K and GuacaMol, sweeping the scalar induces monotone control of logP and consistent QED control across its operating range. The same knob produces opposite structural responses for different targets, arguing against a generic size bias; a negative result under ε-prediction diffusion clarifies where coupling-level alignment is not structurally present.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":98,"slug":130},19,"General","general"]