[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117332-en":3,"doc-seo-117332-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117332,1099514068035,"Ezra","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Information Capacity Regret Bounds for Bandits with Mediator Feedback","This work studies the mediator feedback problem in a multi-armed bandit setting, where each policy corresponds to a probability distribution over outcomes. After selecting a policy, the learner observes a sampled outcome and incurs the loss assigned to that outcome for the current round. The paper introduces policy set capacity as an information-theoretic measure of policy-set complexity, then derives new regret bounds under both adversarial and stochastic regimes using the classical EXP4 algorithm. It also proves nearly-matching lower bounds for selected policy families, extends to varying per-round policies (bandits with expert advice), and shows limitations for exploiting policy similarity in linear bandit feedback, plus a full-information bound via information radius.","POLITECNICO DI TORINO Repository ISTITUZIONALE  \nInformation Capacity Regret Bounds for Bandits with Mediator Feedback  \nOriginal  \nInformation Capacity Regret Bounds for Bandits with Mediator Feedback / Eldowa, Khaled; Cesa-Bianchi, Nicolò; Maria Metelli, Alberto; Restelli, Marcello. -In: JOURNAL OF MACHINE LEARNING RESEARCH. -ISSN 1533-7928. - 25:(2024), pp. 1-36.  \nAvailability:  \nThis version is available at: 11583/2994771 since: 2024-11-25T18:59:03Z  \nPublisher:  \nJournal of Machine Learning Research  \nPublished DOI:  \nTerms of use:  \nThis article is made available under terms and conditions as specified in the corresponding bibliographic description in the repository  \nPublisher copyright  \n(Article begins on next page)  \n17 February 2025  \nInformation Capacity Regret Bounds for Bandits with Mediator Feedback  \nKhaled Eldowa [khaled.eldowa@unimi.it](khaled.eldowa@unimi.it)  \nUniversit􀀒a degli Studi di Milano Milano, 20133, Italy  \nNicol􀀒o Cesa-Bianchi [nicolo.cesa-bianchi@unimi.it](nicolo.cesa-bianchi@unimi.it)  \nUniversit􀀒a degli Studi di Milano and Politecnico di Milano Milano, 20133, Italy  \nAlberto Maria Metelli [albertomaria.metelli@polimi.it](albertomaria.metelli@polimi.it)  \nPolitecnico di Milano Milano, 20133, Italy  \nMarcello Restelli [marcello.restelli@polimi.it](marcello.restelli@polimi.it)  \nPolitecnico di Milano Milano, 20133, Italy  \nEditor: Tor Lattimore  \nAbstract  \nThis work addresses the mediator feedback problem, a bandit game where the decision set consists of a number of policies, each associated with a probability distribution over a common space of outcomes. Upon choosing a policy, the learner observes an outcome sampled from its distribution and incurs the loss assigned to this outcome in the present round. We introduce the policy set capacity as an information-theoretic measure for the complexity of the policy set. Adopting the classical EXP4 algorithm, we provide new regret bounds depending on the policy set capacity in both the adversarial and the stochastic settings.  \nFor a selection of policy set families, we prove nearly-matching lower bounds, scaling similarly with the capacity. We also consider the case when the policies' distributions can vary between rounds, thus addressing the related bandits with expert advice problem, which we improve upon its prior results. Additionally, we prove a lower bound showing that exploiting the similarity between the policies is not possible in general under linear bandit feedback. Finally, for a full-information variant, we provide a regret bound scaling with the information radius of the policy set.  \nKeywords: regret minimization, multi-armed bandits, expert advice, information theory, best of both worlds  \n1. Introduction  \nThe framework of multi-armed bandits (MAB) models sequential decision-making problems with partial feedback. Real-world applications of this framework span a wide array of domains and include problems such as dynamic pricing (Misra et al., 2019) and advert placement (Schwartz et al., 2017) . In the classical non-stochastic MAB problem (Auer et al. , 1995), a learner, faced with a 􀀌xed set of actions (also referred to as \\arms\"), repeatedly interacts with the environment in a series of rounds by selecting an action and subsequently  \n􀀍c2024 Khaled Eldowa, Nicol􀀒o Cesa-Bianchi, Alberto Maria Metelli, and Marcello Restelli.  \nLicense: CC-BY 4.0, see [https://creativecommons.org/licenses/by/4.0/](https://creativecommons.org/licenses/by/4.0/. Attribution)[. Attribution](https://creativecommons.org/licenses/by/4.0/. Attribution) requirements are provided  \nat [http://jmlr.org/papers/v25/24-0227.html](http://jmlr.org/papers/v25/24-0227.html).  \nEldowa, Cesa-Bianchi, Metelli, and Restelli  \nobserving a numerical loss assigned beforehand to this action. The performance of the learner is measured via the notion of regret, which compares the cumulative loss of the learner with that of the best action in hindsight. The minimax regret for this p","cbCaieK9jzyNN9Ps","https://ap.wps.com/l/cbCaieK9jzyNN9Ps","pdf",458476,1,37,"English","en",105,"# Introduction\n## Multi-armed bandits and regret\n## Information leakage and structured feedback\n## Linear bandits and graph-feedback learning\n# Mediator feedback setting\n## Policies as outcome distributions\n## Loss maps and observations","[{\"question\":\"What is the mediator feedback problem studied in the paper?\",\"answer\":\"It is a bandit game where the learner selects a policy from a set, then observes an outcome sampled from that policy’s distribution and suffers the loss associated with the observed outcome in the current round.\"},{\"question\":\"How does the paper measure the complexity of a policy set?\",\"answer\":\"It introduces policy set capacity as an information-theoretic measure capturing how complex the policy set is.\"},{\"question\":\"What regret guarantees are provided and under which settings?\",\"answer\":\"Using the classical EXP4 algorithm, the paper provides new regret bounds that depend on policy set capacity in both adversarial and stochastic settings, and it also proves nearly-matching lower bounds for chosen policy families.\"}]","Information Capacity Regret Bounds for Bandits with Mediator Feedback | PDF",1785675247,93,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"information-capacity-regret-bounds-for-bandits-with-mediator-feedback","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/information-capacity-regret-bounds-for-bandits-with-mediator-feedback/117332/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the mediator feedback problem studied in the paper?","Question",{"text":75,"@type":76},"It is a bandit game where the learner selects a policy from a set, then observes an outcome sampled from that policy’s distribution and suffers the loss associated with the observed outcome in the current round.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the paper measure the complexity of a policy set?",{"text":80,"@type":76},"It introduces policy set capacity as an information-theoretic measure capturing how complex the policy set is.",{"name":82,"@type":73,"acceptedAnswer":83},"What regret guarantees are provided and under which settings?",{"text":84,"@type":76},"Using the classical EXP4 algorithm, the paper provides new regret bounds that depend on policy set capacity in both adversarial and stochastic settings, and it also proves nearly-matching lower bounds for chosen policy families.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]