[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122372-en":3,"doc-seo-122372-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122372,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Bandit problems with fidelity rewards","The fidelity bandits problem extends the K-armed bandit framework by augmenting each arm’s reward with an additional payoﬀ based on the player’s historical loyalty to that arm. Two loyalty mechanisms are proposed: a loyalty-points model where extra reward depends on the number of previous plays, and a subscription model where it depends on the current length of consecutive selections. Both stochastic and adversarial settings are analyzed, including carefully adjusted regret notions. The paper studies increasing, decreasing, and coupon variants, providing regret bounds when sublinear performance holds and worst-case lower bounds otherwise, together with matching algorithms.","Bandit problems with 􀀌delity rewards  \nG􀀓abor Lugosi [gabor.lugosi@upf.edu](gabor.lugosi@upf.edu)  \nICREA  \nBarcelona, Spain, and  \nDepartment of Economics and Business Pompeu Fabra University  \nBarcelona, Spain, and  \nBarcelona Graduate School of Economics Barcelona, Spain  \nCiara Pike-Burke [c.pikeburke@gmail.com](c.pikeburke@gmail.com)  \nDepartment of Mathematics Imperial College London London, UK  \nPierre-Andr􀀓e Savalle [psavalle@cisco.com](psavalle@cisco.com)  \n[Cisco Systems](Cisco Systems), [Inc](Inc).  \nParis, France  \nEditor: Aurelien Garivier  \nAbstract  \nThe 􀀌delity bandits problem is a variant of the K-armed bandit problem in which the reward of each arm is augmented by a 􀀌delity reward that provides the player with an additional payo􀀋 depending on how `loyal' the player has been to that arm in the past. We propose two models for 􀀌delity. In the loyalty-points model the amount of extra reward depends on the number of times the arm has previously been played. In the subscription model the additional reward depends on the current number of consecutive draws of the arm. We consider both stochastic and adversarial problems. Since single-arm strategies are not always optimal in stochastic problems, the notion of regret in the adversarial setting needs careful adjustment. We introduce three possible notions of regret and investigate which can be bounded sublinearly. We study in detail the special cases of increasing, decreasing and coupon (where the player gets an additional reward after every m plays of an arm) 􀀌delity rewards. For the models which do not necessarily enjoy sublinear regret, we provide a worst case lower bound. For those models which exhibit sublinear regret, we provide algorithms and bound their regret.  \nKeywords: multi-armed bandit problem, 􀀌delity reward, regret minimization  \n1. Introduction  \nConsider the problem of a worker searching for a good restaurant for their daily lunches. They are willing to explore the neighborhood in order to 􀀌nd the best restaurant, but also want to have good lunches. This is well modeled as a bandit problem, where the worker has to balance exploration (trying out new or unfamiliar restaurants), and exploitation (having a nice lunch at a favored place) . However, the overall experience and satisfaction of the worker may not depend solely on the chosen restaurant, but also on whether the worker is a  \n©2023 G􀀓abor Lugosi, Ciara Pike-Burke, Pierre-Andr􀀓e Savalle.  \nLicense: CC-BY 4.0, see [https://creativecommons.org/licenses/by/4.0/](https://creativecommons.org/licenses/by/4.0/. Attribution)[. Attribution](https://creativecommons.org/licenses/by/4.0/. Attribution) requirements are provided  \nat [http://jmlr.org/papers/v24/21-1408.html](http://jmlr.org/papers/v24/21-1408.html).  \nLugosi, Pike-Burke, Savalle  \nregular at this restaurant. Indeed, being loyal often leads to better experiences (e.g., waiters are more friendly, or the customer gets a free meal once in a while) . We call the bonus from this loyalty the 􀀌delity reward. In other situations, 􀀌delity can have the opposite e􀀋ect on the reward. For example, restaurants may o􀀋er free drinks to new customers, or customers may get bored of visiting the same restaurant. In this case the 􀀌delity reward decreases.  \nWe consider multi-armed bandit problems where the reward of each arm is augmented by a 􀀌delity reward that provides an additional payo􀀋 depending on how loyal the player has been to a given arm. We consider two models for the 􀀌delity rewards, namely the loyalty-points model and the subscription model.  \nUnder the loyalty-points model, the 􀀌delity reward is a (possibly arm-speci􀀌c) function of the total number of past plays of the arm. An important feature of these loyalty points is that once collected, they stay in the bank. This corresponds to the common practice of loyalty programs in marketing and retail, where the customer is rewarded for loyal behavior (e.g., the loyalty schemes o􀀋ered by airline companies","cbCaioGyUUgmqub2","https://ap.wps.com/l/cbCaioGyUUgmqub2","pdf",465655,1,44,"English","en",105,"# Introduction\n## Fidelity bandits overview\n## Loyalty-points model\n## Subscription model\n## Regret objectives in stochastic and adversarial settings","[{\"question\":\"What is the fidelity bandits problem?\",\"answer\":\"It is a variant of the K-armed bandit problem where each arm’s reward is augmented by a fidelity reward that depends on how loyal the player has been to that arm in the past.\"},{\"question\":\"How do the loyalty-points and subscription models differ?\",\"answer\":\"In the loyalty-points model, extra reward depends on the total number of past plays and accumulated points stay available. In the subscription model, extra reward depends on the current number of consecutive draws, so restarting an arm resets the subscription.\"},{\"question\":\"Why is regret defined differently in the adversarial setting?\",\"answer\":\"Because single-arm strategies are not always optimal in stochastic problems, the standard regret notion for adversarial bandits needs adjustment. The paper introduces three regret notions and studies which can be bounded sublinearly.\"}]","Bandit problems with fidelity rewards | PDF",1785810289,111,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"bandit-problems-with-fidelity-rewards","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/bandit-problems-with-fidelity-rewards/122372/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the fidelity bandits problem?","Question",{"text":75,"@type":76},"It is a variant of the K-armed bandit problem where each arm’s reward is augmented by a fidelity reward that depends on how loyal the player has been to that arm in the past.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do the loyalty-points and subscription models differ?",{"text":80,"@type":76},"In the loyalty-points model, extra reward depends on the total number of past plays and accumulated points stay available. In the subscription model, extra reward depends on the current number of consecutive draws, so restarting an arm resets the subscription.",{"name":82,"@type":73,"acceptedAnswer":83},"Why is regret defined differently in the adversarial setting?",{"text":84,"@type":76},"Because single-arm strategies are not always optimal in stochastic problems, the standard regret notion for adversarial bandits needs adjustment. The paper introduces three regret notions and studies which can be bounded sublinearly.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]