[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86135-en":3,"doc-seo-86135-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86135,962075114765,"Quinn","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Controlling Motion Transfer in Diffusion Transformers via Attention Heads","Diffusion Transformers (DiTs) deliver high-quality, temporally coherent video generation, yet motion transfer—following a reference motion while satisfying a target prompt—remains difficult because DiTs’ internal motion and structure representations are not well understood. A head-level analysis identifies motion-specialized heads and structure-specialized heads using displacement maps and attention-map entropy. Based on this, HALO performs head-aware, controllable motion transfer without parameter updates, refining motion cues via semantic correspondence guidance and preserving structure through selective feature injection.","Controlling Motion Transfer in Diffusion Transformers via Attention Heads  \narXiv :2607 . 11081v1 [ cs .CV] 13 Jul 2026  \nSunyoung Jung 1 *, Jiwoo Park 1 ,2 *, Yoonseok Choi 1, Kyobin Choo 1,  \nMing-Hsuan Yang3, and Seong Jae Hwang 1†  \n1Yonsei University 2 LG Electronics 3 University of California, Merced  \nProject page: [https://sunyj-hxppy.github.io/halo](https://sunyj-hxppy.github.io/halo)  \nFig. 1: Overview. We present HALO , a head-aware controllable motion transfer framework for video Diffusion Transformers, which identifies motion- and structurespecialized attention heads within the model. Leveraging these findings, HALO generates videos that follow the target prompt while remaining motion-and structure-aligned with reference videos, achieving accurate motion transfer.  \nAbstract. Diffusion Transformers (DiTs) have advanced video generation with high-quality, temporally coherent results. However, extending them to motion transfer, which requires following reference motion while aligning with a target prompt, remains challenging due to limited understanding of motion and structure representations within DiTs.  \nWe analyze video DiTs at the attention-head level and identify distinct heads specialized for motion and spatial structure. Based on this insight, we propose a head-aware controllable motion transfer framework that requires no parameter updates. Our method refines motion cues from motion-specialized heads via semantic correspondence guidance and preserves structure through selective feature injection. This head-level control not only enables accurate motion transfer but also provides an interpretable foundation for controllable video generation with DiTs.  \nKeywords: Motion Transfer · Diffusion Transformers · Attention Heads  \n1 Introduction  \nMotion transfer in video generation synthesizes a video that follows the motion of a reference video while adhering to a target prompt. The primary objectives  \n* Equal contribution. † Corresponding author.  \n2 S. Jung et al.  \nare (1) motion fidelity, ensuring temporal adherence to the reference motion, and (2) structural alignment, maintaining spatial layout of the reference [14,48] . Achieving these goals requires modeling of spatio-temporal dependencies, an area in which recent video Diffusion Transformers (DiTs) [25, 40, 46, 47] have shown strong capability. Given their ability to capture spatial structure and temporal dynamics, DiTs have become a natural choice for motion transfer [7, 14, 35] .  \nExisting DiT-based motion transfer approaches, such as noise warping [7], rotary positional embedding manipulation [14], and cross-frame attention optimization [35], offer varying degrees of motion controllability. However, these methods focus on manipulating motion representations without an understanding of how motion and structure are encoded within DiTs. This lack of understanding is due to the distributed functionality of DiTs, which makes their internal mechanisms challenging to analyze [3, 8, 12] . Consequently, generated videos often exhibit seemingly plausible motion yet inaccurate trajectories or misaligned object structures relative to the reference video. As illustrated in Fig. 1, Go-withthe-flow (GWTF) [7], a state-of-the-art method, exhibits motion deviations, including a car moving straight instead of turning and two stormtroopers with misaligned positions compared to the reference.  \nTherefore, we conduct a detailed analysis of video DiTs, focusing on the attention heads to answer the fundamental question: How are motion and structural cues internally encoded in video DiTs? To identify the heads specialized in modeling motion, we introduce the first head-level analysis based on displacement maps. The displacement map [35] encodes motion as patch-wise coordinate differences between frames, making it effective for analyzing motion properties of heads. For structural cues, we compute the visual token attention-map entropy, which represents the uncertainty in patc","cbCaiqO7to4LJkrB","https://ap.wps.com/l/cbCaiqO7to4LJkrB","pdf",44557368,4,1,38,"English","en",105,"# Introduction\n## Problem: motion transfer in video generation\n## Head-level analysis: motion and structure encoding\n## Proposed method: HALO framework","[{\"question\":\"What problem does the document address?\",\"answer\":\"It addresses motion transfer for video diffusion transformers: generating a video that follows reference motion while adhering to a target prompt. The main challenge is limited understanding of how DiTs encode motion and structure internally.\"},{\"question\":\"How does the method analyze motion and structure inside video DiTs?\",\"answer\":\"It performs attention-head-level analysis using displacement maps to capture motion properties and attention-map entropy to assess structural uncertainty. This enables identifying distinct subsets of motion-specialized and structure-specialized heads.\"},{\"question\":\"What is HALO and how does it achieve controllable motion transfer?\",\"answer\":\"HALO is a head-aware framework that enforces both motion fidelity and structural alignment without parameter updates. It builds inter-frame displacement guidance using motion-specific heads, adds semantic correspondence guidance to align motion with target semantics, and preserves structure via selective feature injection.\"}]",1784208823,96,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"controlling-motion-transfer-in-diffusion-transformers-via-attention-heads","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/controlling-motion-transfer-in-diffusion-transformers-via-attention-heads/86135/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the document address?","Question",{"text":75,"@type":76},"It addresses motion transfer for video diffusion transformers: generating a video that follows reference motion while adhering to a target prompt. The main challenge is limited understanding of how DiTs encode motion and structure internally.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the method analyze motion and structure inside video DiTs?",{"text":80,"@type":76},"It performs attention-head-level analysis using displacement maps to capture motion properties and attention-map entropy to assess structural uncertainty. This enables identifying distinct subsets of motion-specialized and structure-specialized heads.",{"name":82,"@type":73,"acceptedAnswer":83},"What is HALO and how does it achieve controllable motion transfer?",{"text":84,"@type":76},"HALO is a head-aware framework that enforces both motion fidelity and structural alignment without parameter updates. It builds inter-frame displacement guidance using motion-specific heads, adds semantic correspondence guidance to align motion with target semantics, and preserves structure via selective feature injection.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]