[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86091-en":3,"doc-seo-86091-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86091,2336464648746,"Skyler","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","WSqD: A Horizon-Free Learning Rate Schedule for Large Model Training","Standard learning-rate schedules like cosine annealing assume a fixed training horizon, which hinders post hoc horizon extension and continued training from intermediate checkpoints. Warmup-stable-decay (WSD) improves resumption by keeping a long constant-rate phase followed by a short linear cooldown, but its peak tuning still depends on the original horizon. WSqD replaces WSD’s stable phase with a shifted inverse-square-root base while keeping the final linear cooldown. In a stochastic convex setting, WSqD attains minimax-optimal O(1/√T) last-iterate convergence and uses a horizon-independent base, needing the horizon only to start cooldown. Experiments on SlimPajama language-model pretraining show WSqD matches or outperforms tuned WSD across horizons with one shared peak rate.","WSqD: A Horizon-Free Learning Rate Schedule for Large Model Training  \nJianhao Ma∗ University of Pennsylvania  \nYuxin Chen∗ University of Pennsylvania  \narXiv :2607 . 10959v 1 [ cs .LG] 12 Jul 2026  \nJuly 14, 2026  \nAbstract  \nStandard learning rate schedules such as cosine annealing are tied to a fixed training horizon, limiting their ability to accommodate post hoc horizon extension. Warmup-stable-decay (WSD) partially addresses this issue by maintaining a long constant-rate phase before a short linear cooldown, allowing training to resume from a pre-decay checkpoint. However, its peak learning rate is still tuned based on the original training horizon and can become suboptimal when training is extended. Motivated by stochastic convex optimization, we propose WSqD (Warmup with Square-root base and linear Decay), a learning rate schedule that replaces WSD’s constant stable phase with a shifted inverse-square-root base while retaining the final linear cooldown. In the stochastic convex setting, WSqD provably attains the minimax-optimal O(1/ √T) last-iterate convergence rate. Importantly, its base learning rate schedule is horizon-independent, and the training horizon is needed only to determine when to begin the final cooldown. Empirically, on language-model pretraining using the SlimPajama corpus, WSqD matches or outperforms carefully tuned WSD and other baselines across multiple training horizons while reusing a single peak learning rate.  \nKeywords: learning rate schedule; continued training; horizon-independent training; stochastic convex optimization; large model training  \n1 Introduction  \nThe learning rate schedule is a core design choice in large language model (LLM) training. A well-designed schedule can substantially improve training efficiency and stability, whereas a poorly chosen one may slow convergence or lead to degraded final performance. Modern LLM training pipelines are increasingly iterative and multi-stage. Rather than fixing the training horizon in advance, practitioners may extend training when evidence from scaling laws suggests that additional compute is likely to yield further gains (Kaplan et al. , 2020 ; Hoffmann et al. , 2022) . They may also continue training on domain-specific corpora (Gururanganet al. , 2020 ; Ke et al. , 2023 ; Gupta et al. , 2023) or organize training into multiple stages that serve different purposes (Ibrahim et al. , 2024) . Consequently, mid-training has emerged as a distinct phase of LLM development (Mo et al. , 2025) . Across these scenarios, it is desirable to resume training from an existing checkpoint and extend the training horizon with little or no re-tuning and without ad hoc modifications to the learning rate schedule. This raises an important question:  \nHow should learning rate schedules be designed to remain effective under post hoc horizon extension?  \n1.1 Prior approaches  \nThe aforementioned continued training requirement is not well served by cosine annealing, a standard fixedhorizon learning rate schedule used in many modern LLM training pipelines, including GPT-3 (Brown et al. ,  \n∗ Department of Statistics and Data Science, University of Pennsylvania. Email: {jianhaom,[yuxinc}@wharton.upenn.edu](yuxinc}@wharton.upenn.edu).  \n(a) Cosine (b) WSD (c) WSqD  \nlearning rate ηt  \nηmax  \nηmin  \n0  \n0 T/2 T iteration number t  \n0 T/2 T iteration number t  \n0 T/2 T iteration number t  \nFigure 1: Illustration of cosine, WSD, and the proposed WSqD learning rate schedules.  \n2020), LLaMA (Touvron et al. , 2023), and Llama 3 (Grattafiori et al. , 2024), to name a few. Cosine annealing explicitly ties the learning rate trajectory to a pre-specified endpoint by smoothly decaying the learning rate from the peak learning rate ηmax to a small terminal learning rate ηmin over the planned horizon: 1  \nηCosinet = ηmin + ηmax~~ ~~−2~~ ~~ηmin 􀀒 1 + cos 􀀒 πtT􀀓􀀓 , t = 1 , . . . , T, (1)  \nwhere T denotes the training horizon, i.e., the total number of training iterations. While th","cbCaijtbd2Qfa24L","https://ap.wps.com/l/cbCaijtbd2Qfa24L","pdf",1081263,4,1,27,"English","en",105,"# Introduction\n## Prior approaches","[{\"question\":\"What is WSqD and how does it address horizon dependence?\",\"answer\":\"WSqD (Warmup with Square-root base and linear Decay) replaces WSD’s stable phase with a shifted inverse-square-root base while retaining the final linear cooldown. Its base learning-rate schedule is horizon-independent, and the horizon is only used to decide when to begin the final cooldown.\"}]",1784208452,68,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"wsqd-a-horizon-free-learning-rate-schedule-for-large-model-training","",{"@graph":36,"@context":77},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/wsqd-a-horizon-free-learning-rate-schedule-for-large-model-training/86091/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"What is WSqD and how does it address horizon dependence?","Question",{"text":75,"@type":76},"WSqD (Warmup with Square-root base and linear Decay) replaces WSD’s stable phase with a shifted inverse-square-root base while retaining the final linear cooldown. Its base learning-rate schedule is horizon-independent, and the horizon is only used to decide when to begin the final cooldown.","Answer","https://schema.org",{"og:url":52,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":98,"slug":130},19,"General","general"]