[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-124573-en":3,"doc-seo-124573-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},124573,1099513958607,"Jiven","https://ap-avatar.wpscdn.com/avatar/100002390cf8733938c?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778829742770036399",8,"Research & Report","Best Arm Identiﬁcation for Stochastic Rising Bandits","Stochastic Rising Bandits studies a bandit setting where the expected rewards of arms grow each time they are selected, representing learning entities whose performance improves over time. The paper targets the Best Arm Identification problem with a fixed round budget, aiming to recommend the arm with the largest expected reward at the end of learning. Two algorithms, R-UCBE and phase-based R-SR, are proposed and their success probabilities are guaranteed, then validated through experiments on synthetic and realistic environments.","Best Arm Identiﬁcation for Stochastic Rising Bandits  \nMarco Mussi 1 Alessandro Montenegro 1 Francesco Trov 1 Marcello Restelli 1 Alberto Maria Metelli 1  \narXiv :2302 .07510v1 [ cs .LG] 15 Feb 2023  \nAbstract  \nStochastic Rising Bandits is a setting in which the values of the expected rewards of the available options increase every time they are selected. This framework models a wide range of scenarios in which the available options are learning entities whose performance improves over time. In this paper, we focus on the Best Arm Identiﬁcation (BAI) problem for the stochastic rested rising bandits. In this scenario, we are asked, given a ﬁxed budget of rounds, to provide a recommendation about the best option at the end of the selection process. We propose two algorithms to tackle the above-mentioned setting, namely R-UCBE, which resorts to a UCB-like approach, and R-SR, which employs a successive reject procedure. We show that they provide guarantees on the probability of properly identifying the optimal option at the end of the learning process. Finally, we numerically validate the proposed algorithms in synthetic and realistic environments and compare them with the currently available BAI strategies.  \n1. Introduction  \nMulti-Armed Bandits (MAB, Lattimore & Szepesvri, 2020a) are a well-known framework that effectively solves learning problems requiring sequential decisions. In the classical MAB formulation, we assume to have a ﬁnite number of options whose selection provides a noisy reward. Given a ﬁnite time horizon, the learner chooses a single option, a.k.a. arm, at each round and observes the corresponding reward. In this work, we focus on Stochastic Rising Bandits (SRB), a speciﬁc instance of the MAB framework in which the expected reward of an arm increases according to the number of times it has been pulled. The task of performing online learning in such a scenario has been recently analyzed from a regret minimization perspective in Metelli et al.(2022) . The authors provide no-regret algorithms for the  \n1Politecnico di Milano, Milan, Italy. Correspondence to: Marco Mussi \u003C[marco.mussi@polimi.it](marco.mussi@polimi.it) >.  \nPreprint. Under Review.  \nCopyright 2023 by the author(s) .  \nSRB setting in both the rested and restless cases.  \nThe SRB setting models several real-world scenarios presenting arms improving their performance over time. A classic scenario that can be modeled through this setting (in the rested ﬂavor) is the so-called Combined Algorithm Selection and Hyperparameter optimization (CASH, Thornton et al., 2013 ; Kotthoff et al., 2017 ; Erickson et al., 2020 ; Liet al., 2020 ; Zller & Huber, 2021), a problem of paramount importance in Automated Machine Learning (AutoML, Yao et al., 2018) . Every arm in CASH represents an algorithm that performs hyperparameter optimization, whose mean reward increases when pulled. A pull of an arm represents a unit of time/computation in which we improve (on average) the hyperparameter selection for the corresponding algorithm by performing a training step. This problem has been handled in a bandit Best Arm Identiﬁcation (BAI) fashion in the works by Li et al. (2020) and Cella et al. (2021) . The former handles the problem by considering rising rested bandits whose arm rewards are deterministic, failing to represent the intrinsic uncertain nature of such processes. Instead, the latter models the problem using bandits with stochastic decreasing losses. However, the major limitation of such a work is the assumption that the expected reward evolves according to a known parametric functional class, whose parameters have to be learned.  \nOriginal Contributions In this work, we address the design of algorithms to solve the BAI task in the SRB setting when a ﬁxed budget is provided.1 More speciﬁcally, we are interested in algorithms able to recommend the arm providing the largest expected reward at the end of the learning process (i.e., when the time budget is over) while ","cbCaivLQQIVJeKZa","https://ap.wps.com/l/cbCaivLQQIVJeKZa","pdf",802310,1,23,"English","en",105,"# Introduction\n## Problem setting and motivation\n## Original contributions\n## Paper structure","[{\"question\":\"What is the Stochastic Rising Bandits (SRB) setting?\",\"answer\":\"SRB is a bandit framework where each arm’s expected reward increases with the number of times it is pulled, modeling performance improvement over time.\"},{\"question\":\"What does the Best Arm Identification (BAI) task require in this paper?\",\"answer\":\"With a fixed budget of rounds, the learner recommends the arm with the highest expected reward at the end of the selection process, despite observing noisy rewards.\"},{\"question\":\"Which two algorithms are proposed, and how do they differ?\",\"answer\":\"The paper proposes R-UCBE, a UCB-like approach for optimism under uncertainty, and R-SR, a successive-reject, phase-based method that removes the worst remaining arm each phase.\"}]","Best Arm Identiﬁcation for Stochastic Rising Bandits | PDF",1785893048,58,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"best-arm-identification-for-stochastic-rising-bandits","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/best-arm-identification-for-stochastic-rising-bandits/124573/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the Stochastic Rising Bandits (SRB) setting?","Question",{"text":75,"@type":76},"SRB is a bandit framework where each arm’s expected reward increases with the number of times it is pulled, modeling performance improvement over time.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does the Best Arm Identification (BAI) task require in this paper?",{"text":80,"@type":76},"With a fixed budget of rounds, the learner recommends the arm with the highest expected reward at the end of the selection process, despite observing noisy rewards.",{"name":82,"@type":73,"acceptedAnswer":83},"Which two algorithms are proposed, and how do they differ?",{"text":84,"@type":76},"The paper proposes R-UCBE, a UCB-like approach for optimism under uncertainty, and R-SR, a successive-reject, phase-based method that removes the worst remaining arm each phase.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]