[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82228-en":3,"doc-seo-82228-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82228,962075114765,"Quinn","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Understanding Schedule-Free Methods in Nonconvex Optimization: Rate Guarantees and Escaping Saddles","Schedule-Free methods are designed to reduce the effort of engineering and tuning learning-rate schedulers while retaining, or even improving upon, the performance of optimizers paired with tuned schedules. Existing results were mainly empirical; rigorous nonconvex convergence theory was largely missing. This paper proves worst-case rate guarantees for Schedule-Free gradient descent and Schedule-Free stochastic gradient descent on smooth, possibly nonconvex objectives using a Lyapunov analysis from the associated continuous-time ODE, and establishes strict-saddle avoidance via an arbitrarily small one-time perturbation.","Understanding Schedule-Free Methods in Nonconvex Optimization: Rate Guarantees and Escaping Saddles  \nJiseok Chae  \nNatural Science Research Institute KAIST Daejeon, Republic of Korea [jsch@kaist.ac.kr](jsch@kaist.ac.kr)  \nDonghwan Kim  \nDepartment of Mathematical Sciences KAIST Daejeon, Republic of Korea [donghwankim@kaist.ac.kr](donghwankim@kaist.ac.kr)  \narXiv :2607 .09 167v 1 [ cs .LG] 10 Jul 2026  \nAbstract  \nSchedule-Free methods have attracted growing interest for alleviating the burden of designing and tuning a learning rate scheduler, while matching and sometimes even outperforming optimizers with tuned schedulers. Despite their strong empirical results, their convergence theory in nonconvex optimization, where modern machine learning objectives typically arise, has remained largely unexplored. In this paper, we provide worst-case analyses of Schedule-Free gradient descent and Schedule-Free stochastic gradient descent, in their standard form and without auxiliary modifications or restrictive conditions, for smooth but possibly nonconvex objectives. Based on a Lyapunov analysis derived from the continuous-time limiting ordinary differential equation associated with these methods, we show that Schedule-Free gradient descent and Schedule-Free stochastic gradient descent achieve the optimal worst-case convergence rates attainable among first-order methods. We further formulate Schedule-Free gradient descent as a nonautonomous dynamical system and prove strict-saddle avoidance under an arbitrarily small one-time perturbation. These theoretical results provide a better understanding of the strong performance that Schedule-Free methods demonstrate.  \n1 Introduction  \nThere are many first-order methods that can be used to solve a minimization problem  \nmin f (x), (1)  \nx∈Rd  \nranging from the established gradient descent (GD) and stochastic gradient descent (SGD) methods to those that are now standard in the machine learning community, such as Adam [17] and AdamW [20] . In modern machine learning tasks, such first-order methods are typically used together with a learning rate scheduler, which changes the learning rate during training according to a predefined rule. However, as training dynamics are largely unpredictable a priori, the choice and design of the learning rate scheduler have traditionally relied heavily on ad hoc approaches and heuristics. Schedule-Free [9] is an optimization scheme that aims to eliminate the need for handcrafted learning rate schedules, which are often difficult to design and tune. The update rule of the Schedule-Free method, which combines the techniques of iterate averaging and a momentum-like interpolation with a given base optimizer, is defined as  \nyk = (1 − β)zk + βxk (2a)  \nzk+1 = zk − γkgk (2b)  \nxk+1 = (1 − ck+1)xk + ck+1zk+1 (2c)  \nPreprint.  \nwith initial points x0 = z0 , where β ∈ [0 , 1] is a fixed constant, {γk }k≥0 is the sequence of learning rates which is often set to be a constant sequence, {ck+1}k≥0 is a predefined sequence of averaging rates such that ck+1 ∈ [0 , 1], and gk is the update direction produced by the base optimizer at yk. For example, Schedule-Free stochastic gradient descent (SF-SGD) uses gk = ∇f(yk ,ζk) for some random variable ζk accounting for the stochasticity in the gradient evaluation, as in standard SGD.  \nA notable interpretation of the Schedule-Free scheme is to understand it as an interpolation between two well-known averaging schemes. When β = 0, under the standard choice ck+1 = k1 , Schedule-Free reduces to Polyak–Ruppert averaging [27, 28], where xk becomes the uniform average over the computed iterates y0 , . . . , yk. On the other extreme, when β = 1, Schedule-Free reduces to a scheme known as primal averaging [24, 31] . Despite this connection, β = 1 is seldom used in Schedule-Free methods. A few common choices for β in practice are 0.9 and 0.98, following the original experiments conducted by Defazio et al. [9] . In this paper, as the case β = 1 ","cbCail2Vsjqb6nGi","https://ap.wps.com/l/cbCail2Vsjqb6nGi","pdf",768690,1,51,"English","en",105,"# Introduction\n## Our Contributions","[{\"question\":\"What problem do Schedule-Free methods target in optimization?\",\"answer\":\"They aim to eliminate the need for handcrafted learning-rate schedulers by using an update rule that combines iterate averaging with a momentum-like interpolation applied to a base optimizer.\"},{\"question\":\"What theoretical results are provided for Schedule-Free methods in nonconvex optimization?\",\"answer\":\"The paper gives worst-case analyses for Schedule-Free gradient descent and Schedule-Free stochastic gradient descent, proving optimal attainable convergence rates for first-order methods on smooth, possibly nonconvex objectives.\"},{\"question\":\"How does the paper handle saddle points for Schedule-Free gradient descent dynamics?\",\"answer\":\"It formulates Schedule-Free gradient descent as a nonautonomous dynamical system and proves strict-saddle avoidance under an arbitrarily small one-time perturbation.\"}]",1784178989,129,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"understanding-schedule-free-methods-in-nonconvex-optimization-rate-guarantees-and-escaping-saddles","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/understanding-schedule-free-methods-in-nonconvex-optimization-rate-guarantees-and-escaping-saddles/82228/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem do Schedule-Free methods target in optimization?","Question",{"text":75,"@type":76},"They aim to eliminate the need for handcrafted learning-rate schedulers by using an update rule that combines iterate averaging with a momentum-like interpolation applied to a base optimizer.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What theoretical results are provided for Schedule-Free methods in nonconvex optimization?",{"text":80,"@type":76},"The paper gives worst-case analyses for Schedule-Free gradient descent and Schedule-Free stochastic gradient descent, proving optimal attainable convergence rates for first-order methods on smooth, possibly nonconvex objectives.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the paper handle saddle points for Schedule-Free gradient descent dynamics?",{"text":84,"@type":76},"It formulates Schedule-Free gradient descent as a nonautonomous dynamical system and proves strict-saddle avoidance under an arbitrarily small one-time perturbation.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]