[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83389-en":3,"doc-seo-83389-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},83389,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","Systematic Evaluation of Learning Rate Scheduling Strategies Across Heterogeneous Architectures","Learning rate (LR) scheduling selection is essential for neural network training, yet manual choice is costly and often non-exhaustive. This study systematically measures how scheduler design affects classification accuracy across heterogeneous architectures, treating the scheduler as the primary varying factor. Thirty representative convolutional and transformer architectures are trained on CIFAR-10 using automated source-code injection to apply 25 scheduler configurations within nine PyTorch families, producing 3,938 evaluated model variants. CosineAnnealingWarmRestarts and CyclicLR outperform basic decay strategies. The resulting accuracy landscape, released through the LEMUR nn-dataset, supports principled scheduler selection.","Systematic Evaluation of Learning Rate Scheduling Strategies Across  \nHeterogeneous Architectures  \nHafsa Mateen*, Radu Timofte, Dmitry Ignatov  \nComputer Vision Lab, CAIDAS & IFI, University ofW¨urzburg, Germany  \narXiv :2607 .085 1 1v 1 [ cs .LG] 9 Jul 2026  \nAbstract  \nChoosing a learning rate scheduling strategy is critical to neural network training, but manual selection is costly and rarely exhaustive. While classical AutoML approaches often treat the scheduler as a secondary hyperparameter, we systematically investigate its impact on classification accuracy across a diverse pool of architectures.  \nWe evaluated 30 representative architectures from convolutional and transformer families within the LEMUR neural network dataset. Through automated source-code injection, we applied 25 scheduler configurations across nine PyTorch families, evaluating a total of 3,938 model vari ants on CIFAR-10. Our best configuration achieved a top-1 accuracy of 86. 45%, with 237 variants exceeding 80% . The results show that the choice of scheduler depends heavily on the architecture: CosineAnnealingWarmRestartsand CyclicLR consistently outperform basic decay strategies. The resulting accuracy landscape, contributed to the LEMUR nn-dataset, provides a practical reference for principled scheduler selection.  \n1. Introduction  \nDeep learning has witnessed remarkable progress over the past decade, driven by advances in model architectures, large-scale datasets, and computational infrastructure. Yet the training recipe—the precise combination of optimizer, learning rate schedule, regularization, and augmentation strategy—remains a largely hand-crafted artifact that requires substantial expert knowledge and empirical trial.  \nAmong the components of a training recipe, the learning rate (LR) schedule exerts a particularly strong influence on convergence speed, final accuracy, and generalization. A poorly chosen schedule can cause slow convergence, oscillation around the optimum, or premature saturation, while a well-chosen one can recover several percentage points of accuracy with no change to the model architecture. Despite this significance, most published works report results with a  \n* Corresponding author: [hafsa.mateen@stud-mail.uni-wuerzburg.de](hafsa.mateen@stud-mail.uni-wuerzburg.de)  \nsingle fixed schedule and treat the scheduler as a secondary hyperparameter.  \nClassical hyperparameter optimization (HPO) methods such as Bayesian optimization [6] and random search can in principle explore the scheduling space, but they are expensive, require repeated full training runs, and offer limited interpretability. Systematic grid-based evaluation across a diverse architecture pool offers a complementary and reproducible alternative: by fixing all other training hyperparameters and varying only the scheduler, one can directly attribute accuracy differences to scheduling choices and build a reusable reference landscape.  \nWe present such a study, contributing a large-scale systematic evaluation of learning rate scheduling strategies across 30 heterogeneous neural network architectures evaluated on CIFAR-10 .  \nOur contributions are as follows.  \n• We contribute 3,938 LR-scheduler model variants to the LEMUR nn-dataset [1, 3, 13] by injecting 25 diverse LR scheduler configurations into 30 neural network architectures via automated source-code injection, evaluating all 3,938 variants on CIFAR-10 for five epochs.  \n• We systematically evaluate the impact of scheduler choice on top-1 accuracy, revealing strong architecture-specific preferences: CyclicLR dominates on mobile-optimized and convolutional models, while CosineAnnealingWarmRestarts leads on inception-based architectures.  \n• We provide a detailed analysis of weight decay interactions, individual scheduler variant rankings, and architecture mean accuracy, offering practical guidance for training recipe design.  \n• We report that the best configuration achieves 86.45% top-1 accuracy, with 237 ","cbCaijNZvv5OQmRI","https://ap.wps.com/l/cbCaijNZvv5OQmRI","pdf",481397,2,1,"English","en",105,"# Introduction\n## Learning rate scheduling impact\n## Systematic evaluation vs HPO\n# Related Work\n## Learning Rate Scheduling\n## Automated Hyperparameter Optimization\n# Method Overview\n## Architecture pool and scheduler catalogue\n## Code injection and evaluation pipeline\n# Experimental Setup\n## Training protocol\n# Results and Analysis\n## Architecture-specific scheduler preferences\n# Comparative Analysis\n## Weight decay interactions and rankings\n# Limitations\n# Conclusion","[{\"question\":\"Why is learning rate scheduling considered important in neural network training?\",\"answer\":\"Learning rate schedules strongly affect convergence speed, final accuracy, and generalization. A poorly chosen schedule can slow convergence, oscillate near the optimum, or saturate prematurely, while a well-chosen one can improve accuracy without changing the model architecture.\"},{\"question\":\"How did the study evaluate learning rate schedulers across architectures?\",\"answer\":\"The study evaluated 30 convolutional and transformer architectures on CIFAR-10. Using automated source-code injection, it applied 25 scheduler configurations across nine PyTorch families, yielding 3,938 model variants evaluated for five epochs.\"},{\"question\":\"Which learning rate schedulers performed best and how does performance depend on architecture?\",\"answer\":\"CosineAnnealingWarmRestarts and CyclicLR consistently outperform basic decay strategies. CyclicLR tends to dominate on mobile-optimized and convolutional models, while CosineAnnealingWarmRestarts leads on inception-based architectures.\"}]",1784187168,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"systematic-evaluation-of-learning-rate-scheduling-strategies-across-heterogeneous-architectures","",{"@graph":35,"@context":84},[36,52,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,46,49],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":20},"https://docshare.wps.com/document/","Document",{"item":47,"name":12,"@type":42,"position":48},"https://docshare.wps.com/document/research-report/",3,{"item":50,"name":13,"@type":42,"position":51},"https://docshare.wps.com/document/systematic-evaluation-of-learning-rate-scheduling-strategies-across-heterogeneous-architectures/83389/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":23,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"Why is learning rate scheduling considered important in neural network training?","Question",{"text":74,"@type":75},"Learning rate schedules strongly affect convergence speed, final accuracy, and generalization. A poorly chosen schedule can slow convergence, oscillate near the optimum, or saturate prematurely, while a well-chosen one can improve accuracy without changing the model architecture.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How did the study evaluate learning rate schedulers across architectures?",{"text":79,"@type":75},"The study evaluated 30 convolutional and transformer architectures on CIFAR-10. Using automated source-code injection, it applied 25 scheduler configurations across nine PyTorch families, yielding 3,938 model variants evaluated for five epochs.",{"name":81,"@type":72,"acceptedAnswer":82},"Which learning rate schedulers performed best and how does performance depend on architecture?",{"text":83,"@type":75},"CosineAnnealingWarmRestarts and CyclicLR consistently outperform basic decay strategies. CyclicLR tends to dominate on mobile-optimized and convolutional models, while CosineAnnealingWarmRestarts leads on inception-based architectures.","https://schema.org",{"og:url":50,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":50},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":28,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":28,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]