[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81583-en":3,"doc-seo-81583-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81583,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","Self-Transcendence: Is External Feature Guidance Indispensable for Accelerating Diffusion Transformer Training","Guiding diffusion transformers with external semantic features can speed training, but it also adds dependencies on pretrained external encoders. Self-Transcendence argues that DiTs can guide their own training using only internal feature supervision. Effective internal guidance must be structurally clean for shallow blocks to separate noise from signal, and semantically discriminative for learning meaningful representations. The method first aligns DiT features with clean VAE latents for a short warm-up, then applies classifier-free guidance on intermediate features. Learned internal features then supervise new DiT training from scratch, improving quality and convergence. ","VISUAL COMPUTING LAB  \narXiv :2601 .07773v3 [ cs .CV] 10 Jul 2026  \nPOLYU VCLAB • ECCV 2026  \n Self-transcendence:  \nIs External Feature Guidance Indispensable for Accelerating Diffusion Transformer Training?  \nLingchen Sun★1,2 Rongyuan Wu★1,2 Zhengqiang Zhang 1,2 Ruibin Li 1 Yujing Sun 1,2  \nShuaizheng Liu 1,2 Lei Zhang†1,2  \n1 The Hong Kong Polytechnic University 2 OPPO Research Institute  \n★ Equal contribution. † Corresponding author ([cslzhang@comp.polyu.edu.hk](cslzhang@comp.polyu.edu.hk)).  \nAbstract. Recent works such as REPA have shown that guiding diffusion models with external semantic features (eg., DINO) can significantly accelerate the training of diffusion transformers (DiTs) . However, the use of pretrained external features as guidance signals introduces additional dependencies. We argue that DiTs actually have the power to guide the training of themselves, and propose Self-Transcendence, an effective method that achieves fast convergence using internal feature supervision only. The desired internal guidance features should meet two requirements: structurally clean to help shallow blocks separate noise from signal, and semantically discriminative to help shallow layers learn effective representations. With this consideration, we first align the DiT features with the clean VAE latent features, a native component of latent diffusion, for a short training phase (eg., 40 epochs) to improve their structural representations, then apply the classifier-free guidance to the intermediate features, enhancing their discriminative capability and semantic expressiveness. These enriched internal features, learned entirely within the model, are used as supervision signals to guide a new DiT training from scratch. Compared to existing self-contained methods, our approach achieves a significant performance boost. It can even surpass REPA, which uses the external DINO features as guidance, in both generation quality and convergence speed for both class-to-image and text-to-image generation tasks. Codes and models can be found at [https://github.com/csslc/Self-Transcendence](https://github.com/csslc/Self-Transcendence).  \nKEYWORDS : Diffusion Transformers, Training Acceleration, Internal Guidance, External Guidance  \n 1 Introduction  \nDiffusion models have emerged as a powerful framework for generative learning, achieving remarkable performance across a wide range of tasks, including image generation [1, 2], video synthesis [3–6], and multi-modal applications [7–10] . Despite the great success, training diffusion transformers (DiTs) [11, 12] remains computationally intensive and suffers from slow convergence. Many methods [13–22] have been developed to stabilize the DiT model training and accelerate the convergence process. Recent studies [13, 15, 23] have highlighted the crucial role of meaningful intermediate representations in both improving training efficiency and enhancing generative capability.  \nTo enrich feature representations, several representation learning strategies have been proposed, including masked training [17, 24–26], contrastive learning [15], and representation alignment [13, 14, 27] . Among them, the pioneering work REPA [13] introduces an effective regularization strategy to align DiT features with external vision encoders such as DINO [28], significantly accelerating model training and improving generation performance. However, this success is highly dependent on external networks and introduces additional dependencies.  \nVisual Computing Lab · The Hong Kong Polytechnic University 1 / 19  \nTo eliminate the reliance on external supervision, recent works [15, 18, 19] have explored self-contained alternatives. Dispersive Loss [15] introduces a plug-and-play regularizer that encourages feature dispersion without requiring pre-training or auxiliary data. SRA [19] and LayerSync [18] instead leverage the discriminative features in deeper layers during training to guide the learning of shallower layers. Specificall","cbCaiakMg8gR64rq","https://ap.wps.com/l/cbCaiakMg8gR64rq","pdf",27094973,4,1,19,"English","en",105,"# 1 Introduction\n## Problem motivation: external vs. internal guidance\n## Prior work and limitations\n## Proposed criteria for internal guidance\n## Overview of Self-Transcendence approach","[{\"question\":\"Why can external feature guidance accelerate diffusion transformer training, and what drawback does it introduce?\",\"answer\":\" External semantic features help diffusion transformers learn faster by providing meaningful intermediate signals, but using pretrained external encoders creates additional dependencies.\"},{\"question\":\"What two requirements must internal guidance features satisfy in Self-Transcendence?\",\"answer\":\" They must be structurally clean to help shallow blocks disentangle noise from signal, and semantically discriminative so shallow layers can learn effective representations.\"},{\"question\":\"How does Self-Transcendence train internal guidance signals before starting full DiT training from scratch?\",\"answer\":\" It first aligns DiT features with clean VAE latent features for a short phase (e.g., 40 epochs) to improve structural representations, then enhances intermediate features using classifier-free guidance. The enriched internal features are used as supervision to train a new DiT from scratch.\"}]",1784174514,48,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"self-transcendence-is-external-feature-guidance-indispensable-for-accelerating-diffusion-transformer-training","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/self-transcendence-is-external-feature-guidance-indispensable-for-accelerating-diffusion-transformer-training/81583/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why can external feature guidance accelerate diffusion transformer training, and what drawback does it introduce?","Question",{"text":75,"@type":76},"External semantic features help diffusion transformers learn faster by providing meaningful intermediate signals, but using pretrained external encoders creates additional dependencies.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What two requirements must internal guidance features satisfy in Self-Transcendence?",{"text":80,"@type":76},"They must be structurally clean to help shallow blocks disentangle noise from signal, and semantically discriminative so shallow layers can learn effective representations.",{"name":82,"@type":73,"acceptedAnswer":83},"How does Self-Transcendence train internal guidance signals before starting full DiT training from scratch?",{"text":84,"@type":76},"It first aligns DiT features with clean VAE latent features for a short phase (e.g., 40 epochs) to improve structural representations, then enhances intermediate features using classifier-free guidance. The enriched internal features are used as supervision to train a new DiT from scratch.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},"General","general"]