[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-124930-en":3,"doc-seo-124930-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},124930,5909877438554,"Maeve","https://ap-avatar.wpscdn.com/avatar/5600025385ad2bf12a7?_k=1778553567797529272",8,"Research & Report","Improved Contrastive Divergence Training of Energy-Based Models - Research summary","Contrastive divergence training for energy-based models is widely used but often suffers from unstable optimization dynamics. The work improves stability by analyzing an extra gradient term that is difficult to compute and usually omitted, showing it is numerically significant yet tractable to estimate. The approach supports better robustness and generation quality via data augmentation and multi-scale processing, and it evaluates stability across model architectures. Experiments demonstrate improved performance across benchmarks and tasks including image generation, out-of-distribution detection, and compositional generation.","MIT Open Access Articles  \nImproved Contrastive Divergence Training of Energy-Based Models  \nThe MIT Faculty has made this article openly available. Please share how this access benefits you. Your story matters.  \nCitation: Du , Yilun , Li , Shuang , Tenenbaum , Joshua and Mordatch , Igor. 2021. \"Improved Contrastive Divergence Training of Energy-Based Models.\" INTERNATIONAL CONFERENCE ON MACHINE LEARNING , VOL 139 , 139.  \nPersistent URL: [https://hdl.handle. net/1721.1/150393](https://hdl.handle. net/1721.1/150393)  \nVersion: Final published version: final published article , as it appeared in a journal , conference proceedings , or other formally published context  \nTerms of Use: Article is made available in accordance with the publisher 's policy and may be subject to US copyright law. Please refer to the publisher 's site for terms of use.  \nImproved Contrastive Divergence Training of Energy-Based Models  \nYilun Du 1 Shuang Li 1 Joshua Tenenbaum 1 Igor Mordatch 2  \nAbstract  \nContrastive divergence is a popular method of training energy-based models, but is known to have difﬁculties with training stability. We propose an adaptation to improve contrastive divergence training by scrutinizing a gradient term that is difﬁcult to calculate and is often left out for convenience. We show that this gradient term is numerically signiﬁcant and in practice is important to avoid training instabilities, while being tractable to estimate. We further highlight how data augmentation and multi-scale processing can be used to improve model robustness and generation quality. Finally, we empirically evaluate stability of model architectures and show improved performance on a host of benchmarks and use cases,such as image generation, OOD detection, and compositional generation.  \n1 Introduction  \nEnergy-Based models (EBMs) have received an inﬂux of interest recently and have been applied to realistic image generation (Han et al., 2019 ; Du & Mordatch, 2019), 3D shapes synthesis (Xie et al., 2018b) , out of distribution and adversarial robustness (Lee et al., 2018 ; Du & Mordatch, 2019 ; Grathwohl et al., 2019), compositional generation (Hinton, 1999 ; Du et al., 2020a), memory modeling (Bartunov et al., 2019), text generation (Deng et al., 2020), video generation (Xie et al., 2017), reinforcement learning (Haarnoja et al., 2017 ; Du et al., 2019), continual learning (Li et al., 2020), protein design and folding (Ingraham et al.; Du et al., 2020b) and biologically-plausible training (Scellier & Bengio, 2017) . Contrastive divergence is a popular and elegant procedure for training EBMs proposed by (Hinton, 2002) which lowers the energy of the training data and raises the energy of the sampled confabulations generated by the model. The model confabulations are generated via an MCMC process (commonly Gibbs sampling or Langevin  \n1MIT CSAIL 2 Google Brain. Correspondence to: Yilun Du \u003C[yilundu@mit.edu](yilundu@mit.edu) >.  \nProceedings of the 38 th International Conference on Machine Learning, PMLR 139, 2021 . Copyright 2021 by the author(s) .  \nFigure 1: (Left) 128x128 samples on unconditional CelebA-HQ.(Right) 128x128 samples on unconditional LSUN Bedroom. dynamics), leveraging the extensive body of research on sampling and stochastic optimization. The appeal of contrastive divergence is its simplicity and extensibility. It does not require training additional auxiliary networks (Kim & Bengio, 2016 ; Dai et al., 2019) (which introduce additional tuning and balancing demands), and can be used to compose models zero-shot.  \nDespite these advantages, training EBMs with contrastive divergence has been challenging due to training instabilities. Ensuring training stability required either combinations of spectral normalization and Langevin dynamics gradient clipping (Du & Mordatch, 2019), parameter tuning (Grathwohl et al., 2019), early stopping of MCMC chains (Nijkampet al., 2019b), or avoiding the use of modern deep learning components, such as self","cbCailzMrRDN3yFv","https://ap.wps.com/l/cbCailzMrRDN3yFv","pdf",8612559,1,13,"English","en",105,"# Abstract\n# Introduction\n## Energy-Based Models and Contrastive Divergence\n## Motivation: Training Instabilities and Common Mitigations\n## The Neglected Gradient Term and Its Efficient Estimation","[{\"question\":\"What problem does improved contrastive divergence training address?\",\"answer\":\"It addresses training instabilities commonly observed when using contrastive divergence to train energy-based models.\"},{\"question\":\"What is the key change proposed by the work?\",\"answer\":\"The method scrutinizes a gradient term introduced by changes to the energy function, estimating it efficiently instead of ignoring it.\"},{\"question\":\"How does the approach improve robustness and generation quality?\",\"answer\":\"It leverages data augmentation and multi-scale processing to enhance mixing in MCMC transitions and improve sample diversity and quality.\"}]","Improved Contrastive Divergence Training of Energy-Based Models - Research summary | PDF",1785895444,33,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"improved-contrastive-divergence-training-of-energy-based-models-research-summary","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/improved-contrastive-divergence-training-of-energy-based-models-research-summary/124930/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does improved contrastive divergence training address?","Question",{"text":75,"@type":76},"It addresses training instabilities commonly observed when using contrastive divergence to train energy-based models.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the key change proposed by the work?",{"text":80,"@type":76},"The method scrutinizes a gradient term introduced by changes to the energy function, estimating it efficiently instead of ignoring it.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the approach improve robustness and generation quality?",{"text":84,"@type":76},"It leverages data augmentation and multi-scale processing to enhance mixing in MCMC transitions and improve sample diversity and quality.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]