[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82174-en":3,"doc-seo-82174-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82174,687197207057,"Sage","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Probing Diffusion Denoising Dynamics for Contrastive Representation Learning","Text-to-image diffusion models offer strong generative capability and rich intermediate representations that can benefit discriminative vision tasks. This work studies how denoising dynamics from a pretrained diffusion model can be adapted for discriminative representation learning while keeping generative behavior under parameter-efficient updates. The proposed D3 CL treats noisy latents at different timesteps as stochastic views of the same image, enabling a contrastive objective combined with denoising reconstruction loss. LoRA-updated Stable Diffusion achieves strong ImageNet-1K results and preserves generation quality.","arXiv :2607 .09067v1 [ cs .CV] 10 Jul 2026  \nProbing Diffusion Denoising Dynamics for Contrastive Representation Learning  \nYasong Dai 1,2 , Zeeshan Hayder 1,2 , David Ahmedt-Aristizabal2 , Hongdong Li 1,3  \n1Australian National University, 2 CSIRO Data61, 3Amazon {yasong.dai, zeeshan.hayder, [hongdong.li}@anu.edu.au](hongdong.li}@anu.edu.au)[ ](hongdong.li}@anu.edu.au){[david.ahmedtaristizabal}@data61.csiro.au](david.ahmedtaristizabal}@data61.csiro.au)  \nAbstract  \nText-to-image diffusion models exhibit unprecedented generative capability and contain rich intermediate representations that can be useful for discriminative vision tasks. Motivated by this observation, we study a focused question: how can the denoising dynamics of a pretrained diffusion model be adapted to support discriminative representation learning while preserving its generative behavior under parameter-efﬁcient updates? We present D3 CL as an investigation of this question.  \nOur key observation is that noisy latents at different diffusion timesteps can be interpreted as stochastic views of the same underlying image, enabling a contrastive objective to be coupled with the standard denoising reconstruction loss. This formulation provides a simple way to probe the interaction between generative denoising and discriminative representation learning without training from scratch.  \nTo keep the adaptation lightweight, we apply LoRA updates to a pretrained Stable Diffusion backbone while freezing the original model parameters. D3 CL provides strong empirical evidence that reconstruction and noise-level contrastive objectives can be complementary: on ImageNet-1K, it obtains 80.1% linear-probing accuracy and an FID of 5.56 for 256 × 256 unconditional generation. Additional ablations on the design space suggest that the usefulness of diffusion features depends on where and how denoising states are sampled. These results establish D3 CL as a parameter-efﬁcient adaptation framework for pretrained diffusion models, showing that noise-level contrastive learning can structure denoising representations for discriminative tasks while maintaining generative performance.  \n1 Introduction  \nSelf-supervised representation learning has demonstrated remarkable results in deriving rich, transferable features without additional supervision signals. Contrastive approaches [2, 1] and generative methods [4, 35] have been developed along separate paths to learn robust visual representations. However, recent research [17, 19] suggests that both contrastive and generative paradigms have shared underlying principles in capturing semantic information from unlabeled data.  \nFollowing this idea, several methods [11, 7, 38] have aimed to unify self-supervised learning for both generative and discriminative tasks. However, these methods still encounter notable limitations, particularly in balancing the trade-off between feature robustness for recognition and high-quality generation [4] . Another challenge arises largely from the extensive computational demands. A stateof-the-art model [11], for example, relies on a heavily parameterized ViT-L/16 backbone with over 400M trainable parameters, requiring 1600 epochs of training. This high resource demand limits the practicality of such models in real-world applications. This raises a critical research question in self-supervised representation learning: Can we develop a uniﬁed framework that effectively balances feature robustness and generation quality while being computationally efﬁcient?  \nPreprint.  \nLinear Probing Accuracy (%) ↑  \n80  \n75  \n70  \n65  \nN/A  \n\n|  |  |  |  |  |  |  | D3 CL( |\n| --- | --- | --- | --- | --- | --- | --- | --- |\n|  | DINO\u003Cbr>DiffF | eed |  |  |  |  | MAGE |\n|  | SimC | LR |  |  |  |  |  |\n| M |  |  |  |  |  | GIVT |  |\n|  |  | Bi | gBiGAN\u003Cbr> |  |  |  |  |\n|  |  | M |  |  | GIT | ICGAN |  |\n\niBOT  \nAE  \nAD  \nMask  \nours)  \nN/A 30 25 20 15 10 5 0 FID (Unconditional Generation) ↓  \nFigure 1: D3 CL balances accuracy and","cbCaicPEUx50kzHD","https://ap.wps.com/l/cbCaicPEUx50kzHD","pdf",5568416,2,1,14,"English","en",105,"# Abstract\n# Introduction\n## Background: self-supervised representation learning\n## Motivation: unify generative and discriminative learning\n## Challenge: robustness–generation trade-off and compute cost\n## Proposed method: D3 CL","[{\"question\":\"What problem does D3 CL address in diffusion-based representation learning?\",\"answer\":\"D3 CL focuses on adapting the denoising dynamics of a pretrained diffusion model for discriminative representation learning while preserving its generative behavior under parameter-efficient updates.\"},{\"question\":\"How does D3 CL connect diffusion denoising with contrastive learning?\",\"answer\":\"D3 CL interprets noisy latents from different diffusion timesteps as stochastic views of the same underlying image, then couples a contrastive objective with the standard denoising reconstruction loss.\"},{\"question\":\"Why use LoRA in this framework?\",\"answer\":\"LoRA enables lightweight parameter-efficient adaptation by updating a pretrained Stable Diffusion backbone while freezing the original model parameters, reducing training overhead.\"}]",1784178592,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"probing-diffusion-denoising-dynamics-for-contrastive-representation-learning","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/probing-diffusion-denoising-dynamics-for-contrastive-representation-learning/82174/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does D3 CL address in diffusion-based representation learning?","Question",{"text":75,"@type":76},"D3 CL focuses on adapting the denoising dynamics of a pretrained diffusion model for discriminative representation learning while preserving its generative behavior under parameter-efficient updates.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does D3 CL connect diffusion denoising with contrastive learning?",{"text":80,"@type":76},"D3 CL interprets noisy latents from different diffusion timesteps as stochastic views of the same underlying image, then couples a contrastive objective with the standard denoising reconstruction loss.",{"name":82,"@type":73,"acceptedAnswer":83},"Why use LoRA in this framework?",{"text":84,"@type":76},"LoRA enables lightweight parameter-efficient adaptation by updating a pretrained Stable Diffusion backbone while freezing the original model parameters, reducing training overhead.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]