[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83891-en":3,"doc-seo-83891-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83891,8796095461564,"Liam","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Consistent and Editable: A Balanced Framework for Text-Guided Video Editing","Text-guided video editing based on diffusion models faces a persistent tension between temporal consistency and editability, since improvements in one often degrade the other. EquiEdit introduces a consistency-and-editability framework that coordinates both objectives. A temporal Mamba module with tailored temporal-aware scanning fuses edited frames along multiple designed directions to strengthen inter-frame consistency. For editability, a spectral-transformation noise injection strategy improves sampling flexibility while preserving latent structure via Fourier transform, supporting fidelity to the input video. Extensive experiments validate consistency, editability, and overall video fidelity.","Consistent and Editable: A Balanced Framework for Text-Guided Video Editing  \nTao Jin 1 * , Li Xiao 1†  \n1 University of Science and Technology of China  \narXiv :2607 .05056v 1 [ cs .CV] 6 Jul 2026  \nAbstract  \nRecently, diffusion models have achieved considerable success in the text-guided video editing domain. However, existing works often struggle to balance the trade-off between temporal consistency and editability in video editing, with consistency and editability typically being inversely related. To address this, we propose a high-quality video editing framework enhanced for consistency and editability, named EquiEdit, which improves coordinatively the temporal consistency and editability of the edited videos while achieving a balance between the two. In terms of temporal consistency, the proposed temporal Mamba module with a tailored temporal-aware scanning scans fused video sequences following four designed directions, effectively enhancing the inter-frame consistency of edited videos. For editability, we design a noise injection strategy based on the spectral transformation to increase editing flexibility, where the Fourier transform is used to preserve the hidden structure in the initial latent noise used for editing, ensuring inter-frame consistency of the edited video and fidelity to the input video. Extensive qualitative and quantitative experiments demonstrate the effectiveness of our method in terms of temporal consistency and editability, as well as its great fidelity to the input video itself.  \nIntroduction  \nAs a cutting-edge generative model in the AIGC wave, diffusion models have sparked rapid development in multiple tasks related to images, including image generation [16, 17, 22], image translation [18, 26], super-resolution [32, 38], and image editing [1, 23, 24] . In image editing tasks guided by text prompts, diffusion-based methods [8, 31] achieve diverse and high-quality results due to their powerful controllability, stability, and amazing realism. This natural language-guided approach not only opens up a new paradigm for image editing, but also provides the possibility for ordinary users, even those without computer expertise, to freely engage in image editing. Success in text-to-image (T2I) editing paves the way for text-to-video (T2V) development.  \nIn the era of short video, compared with static text and images, videos can present dynamic information and are the  \n* First author. E-mail: [jt0618@mail.ustc.edu.cn](jt0618@mail.ustc.edu.cn)  \n†Corresponding author. E-mail: [xiaoli11@ustc.edu.cn](xiaoli11@ustc.edu.cn)  \nA car driving on the dirt road.  \nA car driving on the dirt road, starry night style of Van Gogh.  \n\n| (a) Temporal Inconsistency and Infidelity |  | (b) Lack of Effective Editing |\n| --- | --- | --- |\n\n|  |\n| --- |\n| (c) Our Results |\n\nFig. 1: Illustration of temporal inconsistency and infidelity to the input video (a) and lack of effective editing (b) . For comparison, we present our results in (c) .  \ndominant force on the Internet [42] . Therefore, video editing is of great significance to the media and entertainment industry. However, unlike static image editing, maintaining coherence and temporal consistency between output video frames is one of the main challenges faced by video editing. To solve it, some methods [4, 7, 36] of training a T2V model on large-scale text-video pairs datasets can maintain consistency between output video frames, but they are timeconsuming and computationally expensive, and obtaining large-scale text-video datasets such as WebVid-10M [2] is difficult. On the other hand, videos edited using fine-tuning approaches [46] that introduce temporal modules into pretrained T2I models often exhibit temporal inconsistencies, such as flickering, lack of fidelity to the input video. In Fig. 1 (a), although the edited result from SimDA [41] closely matches the “starry night style of Van Gogh”, it demonstrates inter-frame inconsistencies and infidelity to the inpu","cbCairD2zYAy5voE","https://ap.wps.com/l/cbCairD2zYAy5voE","pdf",2446448,2,1,9,"English","en",105,"# Abstract\n# Introduction\n## Temporal inconsistency vs. lack of effective editing\n## EquiEdit design goals\n## Temporal Mamba for consistency","[{\"question\":\"What problem does EquiEdit address in text-guided video editing?\",\"answer\":\"EquiEdit targets the trade-off between temporal consistency and editability, which many prior methods struggle to balance for coherent edited video frames.\"},{\"question\":\"How does EquiEdit improve temporal consistency?\",\"answer\":\"It uses a temporal Mamba module with temporal-aware scanning that fuses video sequences along multiple designed directions to enhance inter-frame consistency.\"},{\"question\":\"How does EquiEdit improve editability while preserving fidelity?\",\"answer\":\"It proposes a noise injection strategy based on spectral transformation, using Fourier transform to preserve hidden structure in latent noise so edits remain flexible while maintaining fidelity to the input video.\"}]",1784191265,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"consistent-and-editable-a-balanced-framework-for-text-guided-video-editing","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/consistent-and-editable-a-balanced-framework-for-text-guided-video-editing/83891/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does EquiEdit address in text-guided video editing?","Question",{"text":75,"@type":76},"EquiEdit targets the trade-off between temporal consistency and editability, which many prior methods struggle to balance for coherent edited video frames.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does EquiEdit improve temporal consistency?",{"text":80,"@type":76},"It uses a temporal Mamba module with temporal-aware scanning that fuses video sequences along multiple designed directions to enhance inter-frame consistency.",{"name":82,"@type":73,"acceptedAnswer":83},"How does EquiEdit improve editability while preserving fidelity?",{"text":84,"@type":76},"It proposes a noise injection strategy based on spectral transformation, using Fourier transform to preserve hidden structure in latent noise so edits remain flexible while maintaining fidelity to the input video.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]