[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-124609-en":3,"doc-seo-124609-105":30,"detail-sidebar-cat-0-en-105":95},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},124609,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1785132997149421697",8,"Research & Report","Machine learning in and out of equilibrium","Training neural networks with stochastic gradient descent is closely related to diffusion-like natural processes in high-dimensional spaces, such as protein folding or evolutionary dynamics. This study uses a Fokker-Planck framework from statistical physics to analyze long-time stationary behavior. The resulting persistent parameter-space currents link to entropy production along trajectories, whose stationary rate distributions satisfy integral and detailed fluctuation theorems. Numerical tests on nonlinear regression and MNIST confirm universality, while minibatch details alter the effective loss landscape and diffusion matrix. Leveraging this sensitivity, the work introduces an SGLD variant (SGWORLD) with without-replacement minibatching that can reach an equilibrium stationary state for Bayesian posterior sampling.","arXiv :2306 .0352 1v 1 [ cs .LG] 6 Jun 2023  \nMachine learning in and out of equilibrium  \nShishir Adhikari 1,2,3 , Alkan Kabakc¸ ıo˘glu4,7 , Alexander Strang5 , Deniz Yuret6,7 , and Michael Hinczewski 1,*  \n1 Department of Physics, Case Western Reserve University, Cleveland, OH, U.S.A.  \n2 Department of Systems Biology, Harvard Medical School, Boston, MA, U.S.A.  \n3 Department of Data Science, Dana-Farber Cancer Institute, Boston, MA, U.S.A.  \n4 Department of Physics, Koc¸ University, Istanbul, Turkey  \n5 Department of Statistics, University of Chicago  \n6 Department of Computer Engineering, Koc¸ University, Istanbul, Turkey  \n7 Koc¸ University, KUIS AI Center, Istanbul, Turkey  \n* [mxh605@case.edu](mxh605@case.edu)  \nABSTRACT  \nThe algorithms used to train neural networks, like stochastic gradient descent (SGD), have close parallels to natural processes that navigate a high-dimensional parameter space—for example protein folding or evolution. Our study uses a Fokker-Planck approach, adapted from statistical physics, to explore these parallels in a single, unified framework. We focus in particular on the stationary state of the system in the long-time limit, which in conventional SGD is out of equilibrium, exhibiting persistent currents in the space of network parameters. As in its physical analogues, the current is associated with an entropy production rate for any given training trajectory. The stationary distribution of these rates obeys the integral and detailed fluctuation theorems—nonequilibrium generalizations of the second law of thermodynamics. We validate these relations in two numerical examples, a nonlinear regression network and MNIST digit classification. While the fluctuation theorems are universal, there are other aspects of the stationary state that are highly sensitive to the training details. Surprisingly, the effective loss landscape and diffusion matrix that determine the shape of the stationary distribution vary depending on the simple choice of minibatching done with or without replacement. We can take advantage of this nonequilibrium sensitivity to engineer an equilibrium stationary state for a particular application: sampling from a posterior distribution of network weights in Bayesian machine learning. We propose a new variation of stochastic gradient Langevin dynamics (SGLD) that harnesses without replacement minibatching. In an example system where the posterior is exactly known, this SGWORLD algorithm outperforms SGLD, converging to the posterior orders of magnitude faster as a function of the learning rate.  \n1 Introduction  \nOver the last decade, machine learning based on deep neural networks has profoundly impacted a wide variety of fields, including image recognition, natural language processing, health care, finance, manufacturing, autonomous driving, physics, engineering, structural biology, and others 1–3 . Many of these applications involve training a network: finding model parameters that minimize the discrepancy between the ground truth (known from a set of training data) and the predicted value output by the model. The discrepancy (the so-called loss function) depends on the network parameters, which can number in the billions or higher for the most complex problems. The training process then becomes a search for a minimum in a highly multidimensional landscape defined by the loss function4 . Stochastic gradient descent (SGD), along with a multitude of variant methods derived from it, is one of the most popular algorithms for doing this  \nminimization5 . In SGD each step of the training involves an approximation of the loss function, typically by using a small random subset (minibatch) of the total training data available to calculate the gradient with respect to the network parameters. The parameters are then updated ina downward direction defined by this approximate gradient. The resulting training process can be interpreted as biased diffusion on the \"true\" loss landscape defined by ","cbCaidNkkQ5pbTbP","https://ap.wps.com/l/cbCaidNkkQ5pbTbP","pdf",1979675,1,24,"English","en",105,"# Introduction\n## Stochastic gradient descent as biased diffusion\n## Equilibrium versus nonequilibrium stationarity\n## Fluctuation theorems and entropy production\n## Sensitivity to minibatching and posterior sampling","[{\"question\":\"How does the paper relate SGD to processes in statistical physics?\",\"answer\":\"It interprets SGD training as biased diffusion on a multidimensional loss landscape, drawing parallels to physical and biological processes like protein folding and evolution modeled by equilibrium diffusion.\"},{\"question\":\"What characterizes the stationary state of SGD in the long-time limit?\",\"answer\":\"The stationary state is nonequilibrium and exhibits persistent currents in parameter space rather than zero net flow of probability.\"},{\"question\":\"Which principles are shown to hold universally for these training trajectories?\",\"answer\":\"The stationary distributions of entropy production rates satisfy integral and detailed fluctuation theorems, interpreted as nonequilibrium generalizations of the second law.\"},{\"question\":\"How does minibatching choice affect the stationary distribution, and what is SGWORLD?\",\"answer\":\"Simple differences such as minibatching with or without replacement change the effective loss landscape and diffusion matrix, altering the stationary distribution. SGWORLD is a new stochastic gradient Langevin dynamics variant that leverages without-replacement minibatching to converge to an equilibrium posterior distribution faster in a benchmark where the posterior is known exactly.\"}]","Machine learning in and out of equilibrium | PDF",1785893293,60,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":90,"head_meta":92,"extra_data":94,"updated_unix":28},"machine-learning-in-and-out-of-equilibrium","",{"@graph":36,"@context":89},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/machine-learning-in-and-out-of-equilibrium/124609/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81,85],{"name":72,"@type":73,"acceptedAnswer":74},"How does the paper relate SGD to processes in statistical physics?","Question",{"text":75,"@type":76},"It interprets SGD training as biased diffusion on a multidimensional loss landscape, drawing parallels to physical and biological processes like protein folding and evolution modeled by equilibrium diffusion.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What characterizes the stationary state of SGD in the long-time limit?",{"text":80,"@type":76},"The stationary state is nonequilibrium and exhibits persistent currents in parameter space rather than zero net flow of probability.",{"name":82,"@type":73,"acceptedAnswer":83},"Which principles are shown to hold universally for these training trajectories?",{"text":84,"@type":76},"The stationary distributions of entropy production rates satisfy integral and detailed fluctuation theorems, interpreted as nonequilibrium generalizations of the second law.",{"name":86,"@type":73,"acceptedAnswer":87},"How does minibatching choice affect the stationary distribution, and what is SGWORLD?",{"text":88,"@type":76},"Simple differences such as minibatching with or without replacement change the effective loss landscape and diffusion matrix, altering the stationary distribution. SGWORLD is a new stochastic gradient Langevin dynamics variant that leverages without-replacement minibatching to converge to an equilibrium posterior distribution faster in a benchmark where the posterior is known exactly.","https://schema.org",{"og:url":52,"og:type":91,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":93,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":96},[97,101,105,109,113,118,123,126,131,134,138],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},"Exam",70,"exam",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":29,"slug":112},5,"Comic","comic",{"id":114,"doc_module":4,"doc_module_name":46,"category_name":115,"show_sort_weight":116,"slug":117},6,"Technology",50,"technology",{"id":119,"doc_module":4,"doc_module_name":46,"category_name":120,"show_sort_weight":121,"slug":122},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":124,"slug":125},30,"research-report",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":129,"slug":130},9,"Religion & Spirituality",20,"religion-spirituality",{"id":129,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":129,"slug":133},"World Cup","world-cup",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":135,"slug":137},10,"Lifestyle","lifestyle",{"id":139,"doc_module":4,"doc_module_name":46,"category_name":140,"show_sort_weight":110,"slug":141},19,"General","general"]