[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-133251-en":3,"doc-seo-133251-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},133251,962084925782,"Chloe Bennett","https://ap-avatar.wpscdn.com/davatar_9964176cb1d06d4a9deccf72a44ae3dc",8,"Research & Report","MASSIVE ACTIVATIONS ARE THE KEY TO LOCAL DETAIL SYNTHESIS IN DIFFUSION TRANSFORMERS - Paper Abstract","Massive Activations (MAs) are investigated in Diffusion Transformers (DiTs), where their role in visual generation remains largely unexplored despite prior findings in LLMs and Vision Transformers. The study shows MAs occur across all spatial tokens and their distribution is modulated by input timestep embeddings. Crucially, MAs strongly support local detail synthesis while minimally affecting overall semantic content. Based on these insights, Detail Guidance (DG) is proposed as a training-free, MA-driven self-guidance method that integrates with CFG to improve local fidelity and prompt alignment across multiple pre-trained DiTs.","MASSIVE ACTIVATIONS ARE THE KEY TO LOCAL DETAIL SYNTHESIS IN DIFFUSION TRANSFORMERS  \nChaofan Gan 1 ,2 Zicheng Zhao 1 Yuanpeng Tu3 Xi Chen3 Ziran Qin 1 Tieyuan Chen 1 Mehrtash Harandi2 Weiyao Lin 1 ∗  \n1 Shanghai Jiao Tong University, 2Monash University, 3The University of Hong Kong  \n\n| \u003Cbr>Detail Guidance |  |\n| --- | --- |\n| | |\n\n\n| Detail Guidance (DG) Enhances CFG Details |\n| --- |\n| \u003Cbr>|\n\nFigure 1: Visual results of our Detail Guidance (DG). Left: DG explicitly enhances fine-grained visual details, yielding high-quality outputs. Right: DG integrates seamlessly with Classifier-Free Guidance (CFG), allowing for further refinement of details.  \nABSTRACT  \nMassive Activations (MAs) are a well-documented phenomenon across Transformer architectures, and prior studies in both LLMs and ViTs have shown that they play a substantial role in shaping model behavior. However, the nature and function of MAs within Diffusion Transformers (DiTs) remain largely unexplored.  \nIn this work, we systematically investigate these activations to elucidate their role in visual generation. We found that these massive activations occur across all spatial tokens, and their distribution is modulated by the input timestep embeddings.  \nImportantly, our investigations further demonstrate that these massive activationsplay a key role in local detail synthesis, while having minimal impact on the overall semantic content of output. Building on these insights, we propose Detail Guidance (DG), a MAs-driven, training-free self-guidance strategy to explicitly enhance local detail fidelity for DiTs. Specifically, DG constructs a degraded “detail-deficient”  \nmodel by disrupting MAs and leverages it to guide the original network toward higher-quality detail synthesis. Our DG can seamlessly integrate with ClassifierFree Guidance (CFG), enabling joint enhancement of detail fidelity and prompt alignment. Extensive experiments demonstrate that our DG consistently improves local detail quality across various pre-trained DiTs (e.g., SD3, SD3.5, and Flux) .  \n1 INTRODUCTION  \nDiffusion models (Rombach et al., 2022; Saharia et al., 2022) have recently achieved remarkable success across a wide range of generative tasks. Among various architectures, the Transformer (Vaswani  \n∗ Corresponding Author. Project page: [https://ganchaofan0000.github.io/DG](https://ganchaofan0000.github.io/DG)  \nSD3  \nFlux  \nFigure 2: Massive Activations in DiTs. The activation magnitudes of internal hidden states from the middle block (k = N/2) and timestep (t = T/2) . We present the average magnitudes over 1,000 text prompts. Massive Activations (MAs) are consistently concentrated in a few fixed dimensions across all image patch tokens. The MA dimensions remain consistent across all layers (see Figure 14) .  \net al., 2017) has emerged as a powerful and versatile backbone for diffusion models (Peebles & Xie, 2023), thanks to its flexibility and scalability. With the increasing availability of large-scale data and computational resources, many large Diffusion Transformers (DiTs) (Peebles & Xie, 2023; Esseret al., 2024) have recently emerged, achieving state-of-the-art performance in both image and video synthesis (Yang et al., 2024b; Hong et al., 2022; Wan et al., 2025) .  \nAlong with the rapid progress of DiTs, recent studies (Sun et al., 2024; Darcet et al., 2024; Ganet al., 2025) have uncovered an interesting phenomenon known as Massive Activations (MAs) in these Transformer-based models, where rare hidden activations exhibit unusually large magnitudes. Specifically, (Sun et al., 2024; Xiao et al., 2024) identifies the massive activations in Large Language Models (LLMs) and demonstrates that they are essential for long-context learning. Similar activation patterns are observed in Vision Transformers (ViTs), where they are utilized to process global semantic information (Darcet et al., 2024) . More recently, several works (Gan et al., 2025; Fanget al., 2025) have reported the presence of mas","cbCaid8tf2kNVI4q","https://ap.wps.com/l/cbCaid8tf2kNVI4q","pdf",18033259,1,29,"English","en",105,"# Abstract\n# Introduction\n## Massive activations in transformer backbones\n## Characteristics and timestep modulation\n## Activation disruption and effects on semantics vs details\n## Proposed Detail Guidance (DG) and integration with CFG","[{\"question\":\"Where do massive activations occur in Diffusion Transformers, and what controls their distribution?\",\"answer\":\"Massive activations occur across all spatial tokens, and their distribution is modulated by the input timestep embeddings.\"},{\"question\":\"What impact do massive activations have on generated images?\",\"answer\":\"Disrupting massive activations preserves overall semantic content but significantly degrades local visual details, indicating a key role in local detail synthesis.\"},{\"question\":\"How does Detail Guidance (DG) work, and how is it used with CFG?\",\"answer\":\"DG constructs a degraded “detail-deficient” model by disrupting MAs, then uses it to guide the original network toward higher-quality detail synthesis; it can seamlessly integrate with Classifier-Free Guidance (CFG) for joint improvement in detail fidelity and prompt alignment.\"}]","MASSIVE ACTIVATIONS ARE THE KEY TO LOCAL DETAIL SYNTHESIS IN DIFFUSION TRANSFORMERS - Paper Abstract | PDF",1787216374,73,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"massive-activations-are-the-key-to-local-detail-synthesis-in-diffusion-transformers-paper-abstract","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/massive-activations-are-the-key-to-local-detail-synthesis-in-diffusion-transformers-paper-abstract/133251/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-22","2026-08-20",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Where do massive activations occur in Diffusion Transformers, and what controls their distribution?","Question",{"text":76,"@type":77},"Massive activations occur across all spatial tokens, and their distribution is modulated by the input timestep embeddings.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What impact do massive activations have on generated images?",{"text":81,"@type":77},"Disrupting massive activations preserves overall semantic content but significantly degrades local visual details, indicating a key role in local detail synthesis.",{"name":83,"@type":74,"acceptedAnswer":84},"How does Detail Guidance (DG) work, and how is it used with CFG?",{"text":85,"@type":77},"DG constructs a degraded “detail-deficient” model by disrupting MAs, then uses it to guide the original network toward higher-quality detail synthesis; it can seamlessly integrate with Classifier-Free Guidance (CFG) for joint improvement in detail fidelity and prompt alignment.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]