[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117509-en":3,"doc-seo-117509-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117509,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","Delayed Feedback in Kernel Bandits","Black-box optimization of unknown functions via expensive, noisy evaluations can be modeled as a kernel-based bandit (Bayesian optimization) where sequential queries drive improvement. Prior studies assume immediate feedback, which breaks in practice for recommendation systems, clinical trials, and hyperparameter tuning. The work studies stochastically delayed feedback and proposes BPE-Delay, proving regret O(√(Γk(T)T)+E[τ]) under mild delay assumptions, matching near-optimal scaling. Simulations validate the theory, highlighting robust performance even for non-smooth kernels.","Delayed Feedback in Kernel Bandits  \nSattar Vakili 1 Danyal Ahmed 1 Alberto Bernacchia 1 Ciara Pike-Burke 2  \nAbstract  \nBlack box optimisation of an unknown function from expensive and noisy evaluations is a ubiquitous problem in machine learning, academic research and industrial production. An abstraction of the problem can be formulated as a kernel based bandit problem (also known as Bayesian optimisation), where a learner aims at optimising a kernelized function through sequential noisy observations. The existing work predominantly assumes feedback is immediately available; an assumption which fails in many real world situations, including recommendation systems, clinical trials and hyperparameter tuning. We consider a kernel bandit problem under stochastically dela˜yed feedback, and propose an algorithm with  \nO ( pΓk (T)T +E[τ]) regret, where T is the number of time steps, Γk (T) is the maximum information gain of the kernel with T observations, and τ is the delay random variable. This represents a significant improve˜ment over the state of the  \nart regret bound of O(Γk (T)√T + E[τ]Γk (T)) reported in Verma et al. (2022) . In particular, for very non-smooth kernels, the information gain grows almost linearly in time, trivializing the existing results. We also validate our theoretical results with simulations.  \n1. Introduction  \nThe kernel bandit problem is a flexible framework which captures the problem of learning to optimise an unknown function through successive input queries. Typically, the game proceeds in rounds where in each round the learner selects an input point to query and then immediately receives a noisy observation of the function at that point. This observation can be used immediately to improve the learners decision of which point to query next. Due to its general-  \n1MediaTek Research 2Imperial College London. Correspondence to: Sattar Vakili \u003C[sattar.vakili@mtkresearch.com](sattar.vakili@mtkresearch.com) > .  \nProceedings of the 40 th International Conference on Machine Learning, Honolulu, Hawaii, USA. PMLR 202, 2023 . Copyright 2023 by the author(s) .  \nity, the kernel bandit problem has become very popular in practice. In particular, it enables us to sequentially learn tooptimise a variety of different functions without needing to know many details about the functional form.  \nHowever, in many settings where we may want to use kernel bandits, we also have to deal with delayed observations. For example, consider using kernel bandits to sequentially learn to select the optimal conditions for a chemical experiment. The chemical reactions may not be instantaneous, but instead take place at some random time in the future. If we start running a sequence of experiments, we can start new experiments before receiving the stochastically delayed feedback from the previous ones. However, in this situation we have to update the conditions for future experiments before receiving all the feedback from previous experiments. Similar situations arise in recommendation systems, clinical trials and hyperparameter tuning, so it is of practical relevance that we design kernel bandit algorithms that are able to deal with delayed feedback. Moreover, a big challenge for existing kernel bandit algorithms is the computational complexity. Generally speaking, in each round t, algorithms for kernel bandits require fitting a kernel model to the t observed data points which can have an O (t3 ) complexity. To reduce the complexity, there has been a recent interest in considering batch versions of kernel bandit algorithms. These algorithms select τ input values using the same model then update the model after receiving all observations in the batch. This corresponds to a delay of at most τ in receiving each observation, and thus can be thought of as an instance of delayed feedback in kernel bandits.  \nIn this paper, we study the kernel bandit problem with stochastically delayed feedback. We propose Batch Pure Exploration with Delay","cbCaiebJ9AWJ4LpO","https://ap.wps.com/l/cbCaiebJ9AWJ4LpO","pdf",1824745,1,14,"English","en",105,"# Introduction\n## Kernel bandit problem and delayed observations\n## Proposed method and regret guarantees\n## Kernel information gain and practical implications\n## Simulation validation","[{\"question\":\"What problem does the paper address in kernel bandits?\",\"answer\":\"It addresses kernel bandit (Bayesian optimization) learning when feedback is stochastically delayed rather than immediately available.\"},{\"question\":\"Why is delayed feedback important in real applications?\",\"answer\":\"Delayed feedback arises in recommendation systems, clinical trials, hyperparameter tuning, and sequential chemical experiments where outcomes are observed later.\"},{\"question\":\"What algorithm does the paper propose, and what regret bound is proved?\",\"answer\":\"It proposes Batch Pure Exploration with Delays (BPE-Delay) and proves regret scaling of O(√(Γk(T)T)+E[τ]) under mild assumptions on the delay distribution.\"}]","Delayed Feedback in Kernel Bandits | PDF",1785676465,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"delayed-feedback-in-kernel-bandits","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/delayed-feedback-in-kernel-bandits/117509/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in kernel bandits?","Question",{"text":75,"@type":76},"It addresses kernel bandit (Bayesian optimization) learning when feedback is stochastically delayed rather than immediately available.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why is delayed feedback important in real applications?",{"text":80,"@type":76},"Delayed feedback arises in recommendation systems, clinical trials, hyperparameter tuning, and sequential chemical experiments where outcomes are observed later.",{"name":82,"@type":73,"acceptedAnswer":83},"What algorithm does the paper propose, and what regret bound is proved?",{"text":84,"@type":76},"It proposes Batch Pure Exploration with Delays (BPE-Delay) and proves regret scaling of O(√(Γk(T)T)+E[τ]) under mild assumptions on the delay distribution.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]