[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117841-en":3,"doc-seo-117841-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117841,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","CoRe Optimizer - An All-in-One Solution for Machine Learning","Optimization algorithms and their hyperparameters strongly influence training speed and the final accuracy of machine learning models. A well-designed optimizer should converge rapidly and smoothly to low loss, require limited computation, and remain broadly applicable across tasks. This work evaluates the continual resilient (CoRe) optimizer against nine other first-order gradient-based algorithms, including Adam and RPROP, on diverse machine learning problems.","arXiv :2307 . 15663v 1 [ cs .LG] 28 Jul 2023  \nCoRe Optimizer: An All-in-One Solution for Machine Learning  \nMarco Eckhoff∗ and Markus Reiher†  \nETH Zurich, Department of Chemistry and Applied Biosciences, Vladimir-Prelog-Weg 2, 8093 Zurich, Switzerland.  \n(Dated: July 28, 2023)  \nThe optimization algorithm and its hyperparameters can significantly affect the training speed and resulting model accuracy in machine learning applications. The wish list for an ideal optimizer includes fast and smooth convergence to low error, low computational demand, and general applicability. Our recently introduced continual resilient (CoRe) optimizer has shown superior performance compared to other state-of-the-art first-order gradient-based optimizers for training lifelong machine learning potentials. In this work we provide an extensive performance comparison of the CoRe optimizer and nine other optimization algorithms including the Adam optimizer and resilient backpropagation (RPROP) for diverse machine learning tasks. We analyze the influence of different hyperparameters and provide generally applicable values. The CoRe optimizer yields best or competitive performance in every investigated application, while only one hyperparameter needs to be changed depending on mini-batch or batch learning.  \nKeywords: Continual Resilient (CoRe) Optimizer, Adam, RPROP, AdaMax, RMSprop, AdaGrad, AdaDelta, NAG, Momentum, SGD  \n1. INTRODUCTION  \nThe optimization algorithm can crucially determine the training speed and final performance of machine learning (ML) models [1, 2] . Training of an ML model implies that a loss function needs to be minimized. The loss function is usually a sum over training data points. Instead of calculating it simultaneously for the full training data set (deterministic or batch learning), a (semi-)randomly chosen subset of the training data is often employed (stochastic or mini-batch learning) . This approach can accelerate the convergence with respect to the total computation time because the loss-function accuracy increase is sub-linear for larger batch sizes. To update the weights of the ML model, first-order gradientbased iterative optimization schemes are dominating the field, since the memory demand and computation time per step of second-order optimizers is often too high. In general, the optimization aims at a loss function’s local minimum as a function of the model’s weights because it is sufficient for most ML applications to find weight values with low loss rather than the global minimum.  \nThe simplest form of stochastic first-order minimization for high-dimensional parameter spaces is stochastic gradient decent (SGD) [3] . In SGD, the negative gradient of the loss function with respect to each weight is multiplied by a constant learning rate and the product is subtracted from the respective weight in each update. The loss function gradient is adapted in stochastic gradient decent with momentum (Momentum) [4] and Nesterov accelerated gradient (NAG) [5, 6] . These methods aim to improve convergence by a momentum in the weight updates, as the gradients are based on stochastic estimates. In a different fashion, adaptive gradient  \n∗ [eckhoffm@ethz.ch](eckhoffm@ethz.ch)[ ](eckhoffm@ethz.ch)† [mreiher@ethz.ch](mreiher@ethz.ch)  \n(AdaGrad) [7], adaptive delta (AdaDelta) [8], and root mean square propagation (RMSprop) [9] apply the ordinary loss function gradient combined with a weightspecific, adapted learning rate. Adaptive moment estimation (Adam) [10], adaptive moment estimation with infinity norm (AdaMax) [10], and our recently developed continual resilient (CoRe) optimizer [11] combine momentum with individually adapted learning rates. In resilient backpropagation (RPROP) [12, 13] only the sign of the loss function gradient is employed with individually adapted learning rates.  \nApart from these optimizers, which are applied in this work, many more optimizers have been developed for ML applications in recent years. Fo","cbCaidJc8fhueuTX","https://ap.wps.com/l/cbCaidJc8fhueuTX","pdf",6095811,1,10,"English","en",105,"# Introduction\n## Optimization and hyperparameter impact\n## First-order optimizers overview\n## Need for broadly applicable optimizers","[{\"question\":\"Why do optimizer algorithms and hyperparameters matter in machine learning?\",\"answer\":\"They directly affect training speed and the achievable model accuracy by controlling how quickly and smoothly the loss is minimized.\"},{\"question\":\"What is the main goal of the CoRe optimizer work?\",\"answer\":\"Provide a comprehensive performance comparison of CoRe with multiple optimization algorithms and identify generally usable hyperparameter values.\"},{\"question\":\"How does the CoRe optimizer compare to other first-order optimizers in the investigated tasks?\",\"answer\":\"It achieves best or competitive performance across all tested applications, with only one hyperparameter needing adjustment depending on mini-batch or batch learning.\"}]","CoRe Optimizer - An All-in-One Solution for Machine Learning | PDF",1785679946,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"core-optimizer-an-all-in-one-solution-for-machine-learning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/core-optimizer-an-all-in-one-solution-for-machine-learning/117841/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do optimizer algorithms and hyperparameters matter in machine learning?","Question",{"text":75,"@type":76},"They directly affect training speed and the achievable model accuracy by controlling how quickly and smoothly the loss is minimized.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the main goal of the CoRe optimizer work?",{"text":80,"@type":76},"Provide a comprehensive performance comparison of CoRe with multiple optimization algorithms and identify generally usable hyperparameter values.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the CoRe optimizer compare to other first-order optimizers in the investigated tasks?",{"text":84,"@type":76},"It achieves best or competitive performance across all tested applications, with only one hyperparameter needing adjustment depending on mini-batch or batch learning.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":21,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]