[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84660-en":3,"doc-seo-84660-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84660,4810365810221,"Aurora","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Attention Dynamics in Diffusion Models Visual Analytics Framework for Human AI Collaboration","Diffusion-based text-to-image models can synthesize structured visuals, yet interpreting how semantic structure emerges and evolves is difficult. The described visual analytics framework examines attention dynamics through step-indexed token-level cross-attention maps, their temporal concentration, and their spatial relationships. The approach integrates quantitative measures with data-driven stage identification in an interactive workflow. Case studies on a structured 60-prompt Stable-Diffusion-class benchmark reveal recurring, interpretable patterns and show how linked temporal and spatial views support more effective human–AI collaboration.","Attention Dynamics in Diffusion Models: A Visual Analytics Framework  \nfor Human–AI Collaboration  \nYiran Xiao* George Legrady†  \nUniversity of California, Santa Barbara  \nSanta Barbara, California, United States  \narXiv :2607 .02563v1 [ cs .CV] 28 Jun 2026  \ngenerated image  \n\n| (b) step-resolved bird attention |  |  |  |  |  |\n| --- | --- | --- | --- | --- | --- |\n|  |  |  |  |  |  |\n| denoising step\u003Cbr>\u003Cbr>six sampled token maps phase ribbon with transition signal |  |  |  |  |  |\n\nattention dynamics  \nFigure 1: Overview of step-resolved attention analysis. For bird 04 (a bird on a branch), the teaser connects the synthesized image to the temporal and spatial evidence used throughout the system. Panel (b) samples bird attention across denoising steps and summarizes the trajectory with a phase-aware temporal ribbon, highlighting how semantic localization stabilizes over generation. Panels (c–f) show the late-step token-competition view: individual bird and branch maps, their shared support min(A, B) , and the signed spatial difference bird−branch.  \nABSTRACT  \nDiffusion-based text-to-image models can synthesize complex and highly structured visual content, yet the emergence and evolution of semantic structure remain difficult to interpret. Many existing workflows rely on aggregated attention or scalar summaries that separate temporal change from image-space evidence. To address this gap, we present a visual analytics framework for exploring attention dynamics in diffusion models: the step-indexed evolution of token-level cross-attention maps, their temporal concentration, and their spatial relationships. Our approach enables structured analysis of attention behavior across generation steps by integrating quantitative measures with data-driven stage identification inan interactive workflow. Case studies on a structured 60-prompt Stable-Diffusion-class benchmark illustrate recurring, interpretable patterns within this setting and show how linked temporal and spatial views facilitate the observation and discussion of generative processes, supporting more effective human–AI collaboration.  \nIndex Terms: Diffusion Models, Visual Analytics, Explainable AI, Human–AI Collaboration, Interactive Systems  \n1 INTRODUCTION  \nText-to-image diffusion models can synthesize coherent scenes from short prompts [7, 24, 22, 33, 26], yet the process by which  \n* e-mail: [yiranxiao@ucsb.edu](yiranxiao@ucsb.edu)[ ](yiranxiao@ucsb.edu)†e-mail: [glegrady@ucsb.edu](glegrady@ucsb.edu)  \nsemantic structure emerges remains difficult to inspect. This opacity matters for both model developers and downstream creators: a generated image may look plausible, but the user still has little evidence about when the model localized an object, when an attribute became attached to it, or whether two prompt tokens competed for the same spatial region. Existing interfaces therefore leave analysts with a familiar gap between impressive outputs and weak processlevel explanations.  \nCross-attention offers a useful window into this gap because each denoising step relates text tokens to spatial latent positions [6, 25, 8] . In this paper, we define attention dynamics as the step-indexed evolution of token-level cross-attention maps during denoising, including changes in concentration, step-to-step movement, and spatial overlap or dominance between prompt tokens. Common inspection workflows either aggregate attention into a final word-level heatmap or show many per-token curves without the image-space evidence needed to interpret them. Final maps hide temporal reorganization, while scalar curves hide where the corresponding attention mass moves. For visual analytics, the important unit is not only a token map, but a linked trajectory: when attention changes, where it changes, and which prompt tokens share or separate space.  \nWe present a visual analytics framework for step-resolved diffusion attention. The system captures token maps using Diffusion Attentive Attribution Map","cbCait4dR1Q5hNR9","https://ap.wps.com/l/cbCait4dR1Q5hNR9","pdf",4461925,1,5,"English","en",105,"# Introduction\n# Related Work\n# Visual Analytics Framework","[{\"question\":\"What problem does the framework address in diffusion text-to-image models?\",\"answer\":\"It addresses the difficulty of inspecting how semantic structure emerges over denoising steps, which existing interfaces often explain only via aggregated attention or scalar summaries.\"},{\"question\":\"How does the framework define attention dynamics?\",\"answer\":\"Attention dynamics is defined as the step-indexed evolution of token-level cross-attention maps during denoising, including changes in concentration, movement between steps, and spatial overlap or dominance among prompt tokens.\"},{\"question\":\"What visual workflow does the system provide for human–AI collaboration?\",\"answer\":\"The interactive system captures step-resolved token maps, lets users identify global phases on a timeline, uses a step cursor to inspect attention, and provides token-pair comparison views through overlap and signed-difference.\"}]",1784197537,13,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"attention-dynamics-in-diffusion-models-visual-analytics-framework-for-human-ai-collaboration","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/attention-dynamics-in-diffusion-models-visual-analytics-framework-for-human-ai-collaboration/84660/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the framework address in diffusion text-to-image models?","Question",{"text":75,"@type":76},"It addresses the difficulty of inspecting how semantic structure emerges over denoising steps, which existing interfaces often explain only via aggregated attention or scalar summaries.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the framework define attention dynamics?",{"text":80,"@type":76},"Attention dynamics is defined as the step-indexed evolution of token-level cross-attention maps during denoising, including changes in concentration, movement between steps, and spatial overlap or dominance among prompt tokens.",{"name":82,"@type":73,"acceptedAnswer":83},"What visual workflow does the system provide for human–AI collaboration?",{"text":84,"@type":76},"The interactive system captures step-resolved token maps, lets users identify global phases on a timeline, uses a step cursor to inspect attention, and provides token-pair comparison views through overlap and signed-difference.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":21,"slug":137},19,"General","general"]