[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83322-en":3,"doc-seo-83322-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83322,1374391974585,"Genevieve","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Vanilla SGD with Momentum Survives Heavy-Tailed Noise: Convergence Analysis without Gradient Clipping or Normalization","Stochastic gradient descent (SGD) is analyzed for its ability to converge under heavy-tailed noise without relying on gradient clipping or gradient normalization. The study refines convergence results for vanilla SGD and provides a comprehensive first analysis for vanilla SGD with momentum across strongly convex, convex, and nonconvex objectives. Under only heavy-tailed noise assumptions (bounded p-th moments), derived rates are shown to be inferior to clipped or normalized SGD, exposing intrinsic limitations of unmodified methods. Experiments on synthetic functions support the theory.","Vanilla SGD with Momentum Survives Heavy-Tailed Noise: Convergence Analysis without Gradient Clipping or Normalization  \nRyusei Yamada 1 * Naoki Sato 1 * Hideaki Iiduka 1  \n1Meiji University, Japan  \n*Equal contribution  \narXiv :2607 .08 104v 1 [ cs .LG] 9 Jul 2026  \nAbstract  \nStochastic gradient descent (SGD) is a cornerstone of modern optimization. While its performance under heavy-tailed noise is often addressed through specialized modifications such as gradient clipping or normalization, we investigate a more fundamental question: how does vanilla SGD, particularly with momentum, perform in the presence of heavytailed noise? In this paper, we refine existing convergence results for vanilla SGD and, more importantly, provide the first comprehensive convergence analysis of vanilla SGD with momentum for strongly convex, convex, and nonconvex objectives, without employing any gradient control mechanisms. Our results demonstrate that the obtained convergence rates are inferior to the optimal rates achieved by clipped or normalized variants of SGD, thereby revealing inherent limitations of vanilla methods under heavy-tailed noise. The theoretical findings are supported by experiments on synthetic functions.  \n1 INTRODUCTION  \nThis paper considers the following optimization problem:  \nmin f (x) := 1  \nx∈Rd n  \nn X fi (x), i=1  \nwhere each fi (i ∈ [n]) is differentiable. We address three fundamental cases where f is strongly convex, convex, or nonconvex. Empirical risk minimization is a paramount objective that lies at the core of machine learning. As primary solvers for this problem, stochastic gradient descent (SGD)[Robbins and Monro, 1951] and its momentum-based variant [Polyak, 1964, Rumelhart et al., 1986] remain the gold standard in contemporary optimization. A standard assump-  \ntion in the convergence analysis of such stochastic algorithms is bounded variance of stochastic noise, specifically, that the second moment of the error between the stochastic and full gradients is bounded. While numerous seminal results [Nemirovski et al., 2009, Rakhlin et al., 2012, Liu et al., 2020] rely on this assumption, recent experimental evidence suggests that stochastic noise in modern deep learning [Simsekli et al., 2019, Battash et al., 2024, Ahn et al., 2024] and reinforcement learning [Garg et al., 2021] often exhibits heavy tails. Consequently, research has shifted focus toward the convergence behavior of algorithms when the bounded variance assumption is violated, specifically, when stochastic noise possesses only bounded p-th moments for p ∈ (1 , 2] .  \nUnder this bounded p-th moment assumption, the convergence of vanilla SGD is known to break down [Zhang et al., 2020] . To mitigate this issue, clipped SGD, which scales the gradient magnitude to stay within a predefined threshold, has emerged as a robust alternative. Clipped SGD and its variants have been proven to converge in both expectation [Zhang et al., 2020] and with high probability [Cutkosky and Mehta, 2021, Liu et al., 2023, Sadiev et al., 2023, Nguyen et al., 2023a,b, Liu et al., 2024] . Notably, these studies often incorporate gradient normalization (see Table 1), and it has been widely considered essential to clip or normalize gradients to ensure convergence when noise variance is unbounded.  \nRecently, several studies have begun to challenge this notion by exploring the convergence of SGD (with momentum) without clipping under bounded p-th moment assumptions. Hübler et al. [2025] demonstrated that normalized SGD converges with high probability even without clipping. Liu and Zhou [2025] were the first to establish that normalized SGD with momentum converges in expectation, achieving rates of O 􀀐T − 3pp12 􀀑 when the tail index p is known, and O 􀀐T − ~~ ~~p−2p1 􀀑 when it is unknown. Most recently, Fatkhullinet al. [2025] and He and Lu [2025] proved that vanilla SGD without any normalization or clipping can indeed converge  \nTable 1: Comparison of assumptions, algorithmic","cbCaim4XbkOtlyQ4","https://ap.wps.com/l/cbCaim4XbkOtlyQ4","pdf",886160,4,1,31,"English","en",105,"# Introduction\n## Problem setting and noise assumptions\n## Related work and motivation\n## Comparison of convergence rates","[{\"question\":\"What problem does the paper study?\",\"answer\":\"The paper studies the convergence of vanilla stochastic gradient descent, especially with momentum, when stochastic gradients are corrupted by heavy-tailed noise. It considers strongly convex, convex, and nonconvex objectives.\"},{\"question\":\"What mechanisms does the paper avoid?\",\"answer\":\"The analysis is performed without gradient clipping and without gradient normalization. The goal is to understand whether momentum can still ensure convergence under heavy-tailed noise without gradient control.\"},{\"question\":\"How do the paper’s convergence rates compare to clipped or normalized SGD?\",\"answer\":\"The derived convergence rates for vanilla (unclipped and unnormalized) SGD are shown to be inferior to the optimal rates achieved by clipped or normalized variants. This indicates inherent limitations of vanilla methods under heavy-tailed noise.\"}]",1784186728,78,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"vanilla-sgd-with-momentum-survives-heavy-tailed-noise-convergence-analysis-without-gradient-clipping-or-normalization","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/vanilla-sgd-with-momentum-survives-heavy-tailed-noise-convergence-analysis-without-gradient-clipping-or-normalization/83322/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper study?","Question",{"text":75,"@type":76},"The paper studies the convergence of vanilla stochastic gradient descent, especially with momentum, when stochastic gradients are corrupted by heavy-tailed noise. It considers strongly convex, convex, and nonconvex objectives.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What mechanisms does the paper avoid?",{"text":80,"@type":76},"The analysis is performed without gradient clipping and without gradient normalization. The goal is to understand whether momentum can still ensure convergence under heavy-tailed noise without gradient control.",{"name":82,"@type":73,"acceptedAnswer":83},"How do the paper’s convergence rates compare to clipped or normalized SGD?",{"text":84,"@type":76},"The derived convergence rates for vanilla (unclipped and unnormalized) SGD are shown to be inferior to the optimal rates achieved by clipped or normalized variants. This indicates inherent limitations of vanilla methods under heavy-tailed noise.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]