[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84813-en":3,"doc-seo-84813-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84813,2336464648322,"Aria","https://ap-avatar.wpscdn.com/avatar/2200025388227c56fec?_k=1778556882303663488",8,"Research & Report","Steering Optimisation Trajectories in Diffusion Representation Learning","The work investigates why diffusion autoencoders can produce comparable image quality while learning substantially different latent representations. Optimization dynamics explain this behaviour through reconstruction-versus-representation-quality curves that cluster into two early training regimes. Models in the reconstruction regime prioritize image fidelity first, whereas disentanglement improves more gradually in the disentanglement regime. The study hypothesizes controllability via shortcut targeting in the diffusion U-Net and controlled early noise-level exposure, then introduces STEERINGDRL to steer training.","arXiv :2607 .053 19v 1 [ cs .CV] 6 Jul 2026  \nSteering Optimisation Trajectories in Diffusion Representation Learning  \nRajat Rasal⋆ , Avinash Kori, Tian Xia, Ben Glocker  \nImperial College London  \n{rrr2417,t.xia,agk21,[b.glocker}@imperial.ac.uk](b.glocker}@imperial.ac.uk)  \nAbstract  \nWe study why diffusion autoencoders can achieve similar image quality while learning substantially different latent structures. We trace this behaviour to optimisation dynamics; we analyse curves of image reconstruction against latent representation quality, revealing trajectories that organise around two distinct regimes early in training. Models in the reconstruction regime prioritise image fidelity early, whereas those in the disentanglement regime improve reconstruction and disentanglement more gradually. We hypothesise that this behaviour can be influenced by targeting shortcut pathways in the diffusion U-Net and controlling early noise-level exposure, thereby shaping the reconstruction-disentanglement trade-off during training. To steer optimisation toward stronger representations, we introduce STEERINGDRL, combining gated residual U-Nets with a simple noise-level exposure curriculum for training. Across disentanglement benchmarks, STEERINGDRL improves representation quality and reduces seed sensitivity. Our method further extends to spatial disentanglement in object-centric learning, improving segmentation quality on synthetic and real-world datasets.  \n1 Introduction  \nA central hypothesis in representation learning is that decomposing observations into disentangled components [2, 71] enables compositional generalisation in a data-efficient manner [61, 16] . Here, disentanglement refers to the separation of invariant factors in observational data [2] . As generative models are increasingly used for foundational tasks, the ability to learn structured representations becomes critical for controllability, transfer learning, and domain adaptation [44, 7, 38, 31] . However, models can generate high-fidelity images whilst representations are poorly aligned with true causal factors [45, 24] . In our work, we find that diffusion autoencoders, with identical objectives and architectures, can produce similar image quality with substantially different latent representations. Locatello et al. [51] prove that disentangled representation learning (DRL) requires appropriate inductive biases or appropriate distributional assumptions on both the model, such as supervision or regularisation, and the data, such as knowledge of the data-generating process. Given the challenges with obtaining ground-truth labels, research efforts have focused on designing inductive biases for unsupervised attribute disentanglement [26] and spatial disentanglement with object-centric learning [52] . While effective, these methods often trade off between representation learning and generative fidelity [5, 3] . Here, diffusion models provide a promising alternative [62, 83, 94], potentially mitigating this trade-off [89, 24] . Our findings suggest that diffusion autoencoders admit a spectrum of representational solutions, and which one is reached depends on optimisation dynamics. By analysing the reconstruction versus representation quality trade-off during training, we motivate inductive biases that can steer diffusion autoencoders towards better latent representations.  \nDiffusion models exhibit a natural hierarchical structure: coarse semantic information is captured at high noise levels, while fine-grained details are recovered at lower noise levels [80, 49] . Cross  \nPreprint.  \nattention provides a mechanism for querying this hierarchy, aligning conditioning signals with spatial representations across noise levels [25, 76] . This inductive bias underpins both text-to-image and object-centric diffusion models [67], where text tokens or object-slots are aligned with regions in pixel space [85, 32] . In particular, Yang et al. [88] demonstrate that cross-attention induces d","cbCaiobmKfRWmYfh","https://ap.wps.com/l/cbCaiobmKfRWmYfh","pdf",9483296,2,1,32,"English","en",105,"# Introduction\n## Core problem and motivation\n## Diffusion hierarchy and attention inductive bias\n## Optimization dynamics and proposed steering method","[{\"question\":\"What causes diffusion autoencoders to learn different latent structures despite similar image quality?\",\"answer\":\"The behaviour is traced to optimization dynamics, where reconstruction-versus-representation-quality trajectories cluster into two distinct regimes early in training.\"},{\"question\":\"How do the two early training regimes differ?\",\"answer\":\"In the reconstruction regime, models prioritize image fidelity early, while in the disentanglement regime models improve reconstruction and disentanglement more gradually.\"},{\"question\":\"What is STEERINGDRL and how does it improve representation learning?\",\"answer\":\"STEERINGDRL combines gated residual U-Nets with a noise-level exposure curriculum to restrict reconstruction shortcuts and shape early training dynamics, improving representation quality and reducing seed sensitivity, including for spatial disentanglement.\"}]",1784198412,81,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"steering-optimisation-trajectories-in-diffusion-representation-learning","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/steering-optimisation-trajectories-in-diffusion-representation-learning/84813/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What causes diffusion autoencoders to learn different latent structures despite similar image quality?","Question",{"text":75,"@type":76},"The behaviour is traced to optimization dynamics, where reconstruction-versus-representation-quality trajectories cluster into two distinct regimes early in training.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do the two early training regimes differ?",{"text":80,"@type":76},"In the reconstruction regime, models prioritize image fidelity early, while in the disentanglement regime models improve reconstruction and disentanglement more gradually.",{"name":82,"@type":73,"acceptedAnswer":83},"What is STEERINGDRL and how does it improve representation learning?",{"text":84,"@type":76},"STEERINGDRL combines gated residual U-Nets with a noise-level exposure curriculum to restrict reconstruction shortcuts and shape early training dynamics, improving representation quality and reducing seed sensitivity, including for spatial disentanglement.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]