[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86044-en":3,"doc-seo-86044-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86044,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","h-Flow Flexible Flow-based Image Editing via Doob’s h-Transform","Editing images with pre-trained text-to-image flow models requires balancing target prompt alignment with preserving source consistency. Existing techniques often use inversion pipelines or heuristic source-to-target trajectory constructions, which may be architecture-dependent and sensitive to hyperparameters. This paper presents h-flow, a training-free flow-based editing framework grounded in Doob’s h-Transform, extending it to deterministic rectified flow via an equivalent SDE with matching marginals. Dedicated h-functions provide closed-form reconstruction guidance and velocity-based semantic editing signals, with orthogonal decomposition enabling controllable trade-offs across diverse scenarios.","h-Flow: Flexible Flow-based Image Editing via Doob’s h-Transform  \nZehui Guo 1, Zhen Wang2, Junwei Shu 1, Yang Li 1,*, Changbo Wang 1,*,  \nand Long Chen2  \n1 East China Normal University, Shanghai, China  \n[51275901098@stu.ecnu.edu.cn](51275901098@stu.ecnu.edu.cn) , [51265901091@stu.ecnu.edu.cn](51265901091@stu.ecnu.edu.cn) ,  \n[yli@cs.ecnu.edu.cn](yli@cs.ecnu.edu.cn) , [cbwang@cs.ecnu.edu.cn](cbwang@cs.ecnu.edu.cn)  \n2 The Hong Kong University of Science and Technology, Hong Kong SAR, China  \n[zhenwang@ust.hk](zhenwang@ust.hk) , [longchen@ust.hk](longchen@ust.hk)  \narXiv :2607 . 10800v1 [ cs .CV] 12 Jul 2026  \nFig. 1: h-Flow enables faithful text-based editing across diverse scenarios.  \nAbstract. Editing images with pre-trained text-to-image flow models typically requires carefully balancing target alignment with the desired prompt and source consistency with the original image. Existing approaches either rely on inversion-based pipelines or heuristic sourceto-target trajectory constructions, which often depend on architecturespecific designs or are sensitive to hyperparameters. In this paper, we propose h-flow, a training-free and theoretically grounded flow-based editing framework. Inspired by Doob’s h-Transform, we reformulate image  \n∗ Corresponding author.  \n2 Z. Guo et al.  \nediting as conditional generation under multiple terminal events corresponding to source consistency and target alignment. We first extend the classical h-Transform from SDE-based models to the deterministic RF framework by constructing an equivalent SDE with identical marginals.  \nWithin this formulation, we design dedicated h-functions for source consistency and target alignment, yielding closed-form reconstruction guidance and velocity-based semantic editing signals. We further introduce a velocity orthogonal decomposition to decouple reconstruction and editing directions, enabling a controllable trade-off between the two objectives.  \nExtensive experiments demonstrate that h-flow achieves effective, robust, and flexible editing across diverse scenarios. The code will be released soon.  \nKeywords: Text-based Image Editing · Rectified Flow Model · Doob’sh-Transform  \n1 Introduction  \nRectified flow (RF) models [21,22] have emerged as a powerful generative paradigm for various visual generation tasks. Benefiting from large-scale pretraining, textto-image RF models [9, 20] have been widely adopted for text-based image editing [19,29], which aims to modify the visual contents of a source image according to a target prompt. The central challenge of this task lies in balancing the two common goals: target alignment and source consistency. Specifically, the synthesized image should faithfully adhere to the target prompt while preserving the editing-irrelevant regions unchanged from the source image.  \nTo achieve this trade-off, early flow-based editing approaches predominantly adopt a two-stage, inversion-based paradigm [18, 29, 34] . In the forward process, the source image is first inverted into a latent variable in the noise distribution, in the backward process, the edited image is generated from this latent under the guidance of the target prompt. Representative works mainly focus on improving inversion accuracy [18, 29] for better source reconstruction, and further inject source conditions (e.g., attention maps [34]) during the forward process to enhance source consistency. However, these strategies are largely heuristic and often rely on architecture-specific analysis, limiting their scalability and general applicability. More recent studies have pioneered inversion-free paradigms [13,17,19], which directly construct transformation trajectories from the source image to the target image, enforcing strong structural consistency. Despite their empirical effectiveness, they remain sensitive to specific hyperparameter settings, such as random seeds and the number of flow steps, which constrain their robustness and generalization in practical scenario","cbCaimWesRL1rQtX","https://ap.wps.com/l/cbCaimWesRL1rQtX","pdf",17528407,4,1,33,"English","en",105,"# Abstract\n# Introduction\n## Balancing target alignment and source consistency\n## Limitations of inversion-based approaches\n## Limits of inversion-free trajectory methods\n## h-Flow: Doob’s h-Transform for rectified flow editing","[{\"question\":\"What core challenge does h-flow address in text-based image editing?\",\"answer\":\"It balances two goals: target alignment to the prompt while preserving source-consistent regions from the input image.\"},{\"question\":\"How does h-flow differ from prior inversion-based and inversion-free editing methods?\",\"answer\":\"h-flow is training-free and provides a unified theoretical formulation inspired by Doob’s h-Transform, avoiding architecture-specific heuristics and hyperparameter sensitivity.\"},{\"question\":\"What mechanisms does h-flow use to control editing and reconstruction trade-offs?\",\"answer\":\"It introduces dedicated h-functions for source consistency and target alignment, then applies a velocity orthogonal decomposition to separate reconstruction and editing directions for a controllable balance.\"}]",1784208048,83,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"h-flow-flexible-flow-based-image-editing-via-doobs-h-transform","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/h-flow-flexible-flow-based-image-editing-via-doobs-h-transform/86044/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What core challenge does h-flow address in text-based image editing?","Question",{"text":75,"@type":76},"It balances two goals: target alignment to the prompt while preserving source-consistent regions from the input image.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does h-flow differ from prior inversion-based and inversion-free editing methods?",{"text":80,"@type":76},"h-flow is training-free and provides a unified theoretical formulation inspired by Doob’s h-Transform, avoiding architecture-specific heuristics and hyperparameter sensitivity.",{"name":82,"@type":73,"acceptedAnswer":83},"What mechanisms does h-flow use to control editing and reconstruction trade-offs?",{"text":84,"@type":76},"It introduces dedicated h-functions for source consistency and target alignment, then applies a velocity orthogonal decomposition to separate reconstruction and editing directions for a controllable balance.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]