[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122370-en":3,"doc-seo-122370-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122370,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Delayed Feedback in Generalised Linear Bandits Revisited","The stochastic generalised linear bandit is a standard framework for sequential decision-making, where many methods guarantee near-optimal regret when rewards arrive immediately. In practice, rewards are frequently delayed and can be observed at unknown future times. This work studies delayed rewards within the generalised linear bandit model and analyzes regret under such feedback lags. It shows that an optimistic adaptation of existing algorithms can recover meaningful regret guarantees, and it is evaluated through experiments on simulated data.","Delayed Feedback in Generalised Linear Bandits Revisited  \nBenjamin Howson  \nImperial College London  \nCiara Pike-Burke  \nImperial College London  \nSarah Filippi  \nImperial College London  \nAbstract  \nThe stochastic generalised linear bandit is a well-understood model for sequential decisionmaking problems, with many algorithms achieving near-optimal regret guarantees under immediate feedback. However, the stringent requirement for immediate rewards is unmet in many real-world applications where the reward is almost always delayed. We study the phenomenon of delayed rewards in generalised linear banditsin a theoretical manner. We show that a natural adaptation of an optimistic algorithm to the delayed feedback setting can achieve regret of Tdn+eld3ayTh, /2disi[τthgn]ei)fi,dwhimcaneetrenslyiEonim[]dndroevoss ththepoexmexe--isting approaches for this setting where the best knWoewvrirefygetur nd waoreticaru(√ltTthd + Eugh ex[τpe] r)-.  \niments on simulated data.  \n1 INTRODUCTION  \nRecently, bandit algorithms have found application in areas from dynamic pricing and healthcare to finance and recommender systems with great success (Misra et al., 2019; Durand et al., 2018; Shen et al., 2015; McInerney et al., 2018) . There are many formulations of bandit problems. One of these is the stochastic generalised linear bandit, which captures a wide class of problems, such as when the rewards are counts, binary values or can take any real-valued number. The generalised linear bandit problem proceeds in rounds, where in each round, a learner must choose from a set of possible actions. After selecting an action, the learner receives feedback from the environment in the form of a reward which stochastically depends on the inner product of the action and some unknown parameter vector. The goal  \nProceedings of the 26th International Conference on Artificial Intelligence and Statistics (AISTATS) 2023, Valencia, Spain. PMLR: Volume 206 . Copyright 2023 by the author(s) .  \nof the learner is to maximise their expected cumulative reward.  \nThere are many provably efficient algorithms for the generalised linear bandit (Filippi et al., 2010; Abbasi-Yadkori et al., 2011; Li et al., 2017; Faury et al., 2020) . Unfortunately, these existing algorithms require immediate feedback from the environment. This strict requirement for immediate rewards often goes unmet in practice. For example, in many recommender systems, the user must provide feedback to the learner while operating on a very different time scales; e.g. the learner can make thousands of recommendations per second, whereas a user may take several minutes to a couple of days to respond to the recommendation, if at all (Chapelle, 2014) . Alternatively, practitioners might want to optimise for a longer-term measure of success (Han and Arndt, 2021), in which case the reward is not observable or even defined immediately. Delayed feedback also arises in clinical trials due to the time-consuming task of obtaining medical feedback and because patients do not respond to their prescribed treatment immediately.  \nIn all the above settings, the reward for any given action returns at an unknown time in the future. Meanwhile, the learner must continue operating in the environment without feedback from many of their past choices. A natural model for this phenomenon is to introduce a random delay between taking action and receiving the reward. However, the delays pose significant theoretical challenges because standard tools for analysing bandit algorithms rely on utilising immediate feedback to reduce the uncertainty in the learner’s estimation. Under delayed feedback, it is unclear how long the learner will have to wait before they gain information about the quality of an action, which hinders their future decision-making abilities.  \nThese challenges have led to the development of algorithms specifically for delayed feedback in generalised linear bandits. However, to the best of our knowledge, these existin","cbCaim6trkLo1Grk","https://ap.wps.com/l/cbCaim6trkLo1Grk","pdf",1454513,1,25,"English","en",105,"# Introduction\n## Related Work","[{\"question\":\"What problem does delayed feedback introduce in generalised linear bandits?\",\"answer\":\"Rewards for chosen actions arrive at unknown future times, so the learner must keep making decisions without feedback from many earlier actions. This breaks standard analysis methods that depend on immediate rewards to reduce uncertainty.\"},{\"question\":\"Why do existing delayed-feedback algorithms perform poorly in this setting?\",\"answer\":\"Many methods require prior knowledge of the expected delay, strong assumptions on the delay distribution, or restrictive assumptions on the action sets. Some also yield regret bounds whose delay impact grows with the time horizon.\"},{\"question\":\"What does the paper propose to handle delayed rewards?\",\"answer\":\"It studies a natural optimistic adaptation of an algorithm to the delayed feedback setting and proves a regret bound that captures how delays affect performance.\"}]","Delayed Feedback in Generalised Linear Bandits Revisited | PDF",1785810280,63,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"delayed-feedback-in-generalised-linear-bandits-revisited","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/delayed-feedback-in-generalised-linear-bandits-revisited/122370/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does delayed feedback introduce in generalised linear bandits?","Question",{"text":75,"@type":76},"Rewards for chosen actions arrive at unknown future times, so the learner must keep making decisions without feedback from many earlier actions. This breaks standard analysis methods that depend on immediate rewards to reduce uncertainty.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why do existing delayed-feedback algorithms perform poorly in this setting?",{"text":80,"@type":76},"Many methods require prior knowledge of the expected delay, strong assumptions on the delay distribution, or restrictive assumptions on the action sets. Some also yield regret bounds whose delay impact grows with the time horizon.",{"name":82,"@type":73,"acceptedAnswer":83},"What does the paper propose to handle delayed rewards?",{"text":84,"@type":76},"It studies a natural optimistic adaptation of an algorithm to the delayed feedback setting and proves a regret bound that captures how delays affect performance.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]