[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119167-en":3,"doc-seo-119167-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119167,8796095461564,"Liam","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Stochastic Langevin Differential Inclusions with Applications to Machine Learning","Stochastic differential equations in Langevin form underpin Bayesian sampling and optimization in machine learning, yet their usual theory assumes a smooth potential so the drift is globally Lipschitz. Many learning setups violate this smoothness, making the drift non-Lipschitz, as seen with robust losses and ReLU-based models. This work studies Langevin-type stochastic differential inclusions using assumptions tailored to machine-learning settings, proving strong existence of solutions and establishing asymptotic minimization of the associated free-energy functional.","arXiv :2206 . 11533v3 [math .OC] 12 May 2024  \nStochastic Langevin Differential Inclusions with Applications  \nto Machine Learning  \nFabio V. Difonzo, Vyacheslav Kungurtsev, and Jakub Mareˇcek  \nAbstract  \nStochastic differential equations of Langevin-diffusion form have received significant attention, thanks to their foundational role in both Bayesian sampling algorithms and optimization in machine learning. In the latter, they serve as a conceptual model of the stochastic gradient flow in training over-parameterized models. However, the literature typically assumes smoothness of the potential, whose gradient is the drift term. Nevertheless, there are many problems for which the potential function is not continuously differentiable, and hence the drift is not Lipschitz continuous everywhere. This is exemplified by robust losses and Rectified Linear Units in regression problems. In this paper, we show some foundational results regarding the flow and asymptotic properties of Langevin-type Stochastic Differential Inclusions under assumptions appropriate to the machine-learning settings. In particular, we show strong existence of the solution, as well as an asymptotic minimization of the canonical free-energy functional.  \n1. Introduction  \nIn this paper, we study the following stochastic differential inclusion,  \ndXt ∈ −F(Xt)dt +√2σ dBt (1)  \nwherein F (x) : Rn ⇒ Rn is a set-valued map. We are particularly interested in the case where F (x) is the Clarke subdifferential of some continuous tame function f (x) . This is motivated by the recent interest in studying Langevin-type diffusions in the context of machine learning applications, both as a scheme for sampling in a Bayesian framework (as spurred by the seminal work Welling and Teh (2011)) and as a model of the trajectory of stochastic gradient descent, with a view to understanding the asymptotic properties of training deep neural networks Hu et al. (2019) . It is typically assumed that F (x) above is Lipschitz and as such its potential f (x) is continuously differentiable. In many problems of relevance, including empirical risk minimization with robust loss (e.g. , l1 or Huber) and neural networks with ReLU activations, this is not the case, and yet there is at least partial empirical evidence suggesting that the long-term behavior of a numerically similar operation is similar in its capacity to generate a stochastic process which minimizes a Free Energy associated with the learning problem.  \nThis paper is organized as follows. In Section 2 we study the functional analytical properties of F(x) as it appears in (1) when it represents a noisy estimate of a subgradient element of an empirical loss function that itself satisfies the conditions of a definable potential, especially as it appears in the context of deep learning applications. Note that the research program undertaken relates to a recent conjecture of Bolte and Pauwels (Bolte and Pauwels, 2021, Remark 12) that suggests the strong convergence of iterates in a stochastic subgradient type sequence to stationary points for this class of potentials. Subsequently,  \nF.V. Difonzo, V. Kungurtsev and J. Mareˇcek  \nin Section 3 we prove that there exists a strong solution to (1), confirming the existence of a trajectory in the general case. In this sense, we extend the work of Leobacher and Sz¨olgyenyi (2017); Leobacher and Steinicke (2022) studying diffusions with discontinuous drift to set-valued drift. Next in Section 4 we prove the correspondence of a Fokker-Planck type equation to modeling the probability law associated with this stochastic process, and show that it asymptotically minimizes a free-energy functional corresponding to the loss function of interest, extending the seminal work of Jordan et al. (1998) which had proven the same result in the case of continuous F (x) . We present some numerical results that confirm the expected asymptotic behavior of (1) in Section 5 and summarize our findingsand their implicati","cbCairQ1cqVMC3FY","https://ap.wps.com/l/cbCairQ1cqVMC3FY","pdf",1260481,1,26,"English","en",105,"# Abstract\n# Introduction\n## Related Work\n# Main Results (1)\n## Strong solution existence\n# Main Results (2)\n## Fokker-Planck correspondence and asymptotic minimization\n# Numerical Results\n# Findings and Implications","[{\"question\":\"Why does the classical Langevin SDE theory not directly apply to many machine-learning problems?\",\"answer\":\"Classical results typically assume the potential is smooth, making the drift globally Lipschitz. Robust losses and ReLU activations produce potentials that are not continuously differentiable, so the drift becomes non-Lipschitz.\"},{\"question\":\"What is the stochastic model studied in this paper?\",\"answer\":\"The paper studies a Langevin-type stochastic differential inclusion where the drift is set-valued, using a set-valued map such as the Clarke subdifferential of a continuous tame function.\"},{\"question\":\"What key properties are proven for the proposed stochastic differential inclusion?\",\"answer\":\"The paper proves strong existence of solutions and shows that the dynamics asymptotically minimize a canonical free-energy functional connected to the learning problem.\"}]","Stochastic Langevin Differential Inclusions with Applications to Machine Learning | PDF",1785722883,66,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"stochastic-langevin-differential-inclusions-with-applications-to-machine-learning","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/stochastic-langevin-differential-inclusions-with-applications-to-machine-learning/119167/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04","2026-08-03",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why does the classical Langevin SDE theory not directly apply to many machine-learning problems?","Question",{"text":76,"@type":77},"Classical results typically assume the potential is smooth, making the drift globally Lipschitz. Robust losses and ReLU activations produce potentials that are not continuously differentiable, so the drift becomes non-Lipschitz.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What is the stochastic model studied in this paper?",{"text":81,"@type":77},"The paper studies a Langevin-type stochastic differential inclusion where the drift is set-valued, using a set-valued map such as the Clarke subdifferential of a continuous tame function.",{"name":83,"@type":74,"acceptedAnswer":84},"What key properties are proven for the proposed stochastic differential inclusion?",{"text":85,"@type":77},"The paper proves strong existence of solutions and shows that the dynamics asymptotically minimize a canonical free-energy functional connected to the learning problem.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]