[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84411-en":3,"doc-seo-84411-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84411,1099513958607,"Jiven","https://ap-avatar.wpscdn.com/avatar/100002390cf8733938c?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778829742770036399",8,"Research & Report","Randomized Confidence Bounds for Stochastic Partial Monitoring","The partial monitoring (PM) framework models sequential learning with incomplete feedback, where each round selects an action while the environment chooses an outcome and the agent receives a signal only partially revealing the (unobserved) outcome. The agent uses these signals to minimize cumulative loss in both contextual and non-contextual stochastic settings with i.i.d. outcomes. The work introduces randomized confidence-bound strategies that extend regret guarantees beyond cases where existing stochastic methods apply, and validates them experimentally. A monitoring use case further motivates adoption in real systems.","Randomized Confidence Bounds for Stochastic Partial Monitoring  \nMaxime Heuillet 1 2 3 4 Ola Ahmad 1 2 Audrey Durand 1 3 4 5  \narXiv :2402 .05002v 3 [ cs .LG] 12 Jul 2026  \nAbstract  \nThe partial monitoring (PM) framework provides a theoretical formulation of sequential learning problems with incomplete feedback. On each round, a learning agent plays an action while the environment simultaneously chooses an outcome.  \nThe agent then observes a feedback signal that is only partially informative about the (unobserved) outcome. The agent leverages the received feedback signals to select actions that minimize the (unobserved) cumulative loss. In contextual PM, the outcomes depend on some side information that is observable by the agent before selecting the action on each round. In this paper, we consider the contextual and non-contextual PM settings with stochastic outcomes. We introduce a new class of strategies based on the randomization of deterministic confidence bounds, that extend regret guarantees to settings where existing stochastic strategies are not applicable. Our experiments show that the proposed RandCBP and RandCBPside⋆ strategies improve state-ofthe-art baselines in PM games. To encourage the adoption of the PM framework, we design a use case on the real-world problem of monitoring the error rate of any deployed classification system.  \n1. Introduction  \nPartial monitoring (Bartk et al., 2014) is a framework tailored for online learning problems with partially informative feedback. A partial monitoring (PM) game is played between a learning agent and the environment over multiple rounds. At a given round, the agent selects an action and the environment simultaneously selects an outcome. The agent then incurs an instant loss and receives a feedback signal that is partially informative about the outcome. The challenge  \n*Equal contribution 1Universit Laval 2Thales Digital Solutions (cortAIx lab) 3Mila-Qubec AI Institute 4Institut Intelligence et Donnes 5Canada CIFAR AI Chair. Correspondence to: Maxime Heuillet \u003C[maxime.heuillet.1@ulaval.ca](maxime.heuillet.1@ulaval.ca) > .  \nProceedings of the 41st International Conference on Machine Learning, Vienna, Austria. PMLR 235, 2024 . Copyright 2024 by the author(s) .  \nis that the agent does not observe the loss. Nonetheless, its goal is to minimize the (unobserved) cumulative loss by carefully balancing between actions associated to informative feedback signals and small-loss actions, which captures the core exploration-exploitation trade-off.  \nThe agent’s performance is measured by the regret, which corresponds to the excess loss associated with the selected action compared to the best action in hindsight. The cumulative regret scales linearly with the horizon T if the agent fails to identify the best action. In this work, we consider the stochastic setting where outcomes are independent and identically distributed (i.i.d.) according to some (unknown) outcome distribution. In this setting, Bartk et al. (2011) classified PM games into four categories based on achievable bounds on the cumulative regret: trivial games (noregr˜et); easy games with poly-logarithmic upper bounds  \nin Θ( √T ) ; hard games with upper bounds in Θ(T2/3) ; and intractable games with lower bounds in Ω(T) . The well-known multi-armed bandit problem (Auer et al., 2002) corresponds to an easy game. Additionally, many problems correspond to hard games, such as learning from costly expert advice (Helmbold & Panizza, 1997), dynamic pricing (Kleinberg et al., 2003), and online monitoring (Ginart et al., 2022) .  \nDeterministic PM strategies such as CBP (Bartk et al., 2012) and PMDMED (Komiyama et al., 2015) have sub-linear regret guarantees on both easy and hard games. However, these are consistently outperformed empirically by stochastic strategies like BPM-Least (Vanchinathan et al., 2014) and TSPM (Tsuchiya et al., 2020), for which regret guarantees are unfortunately limited to easy games.  \nThe context","cbCaityEitHggVrx","https://ap.wps.com/l/cbCaityEitHggVrx","pdf",1007386,1,36,"English","en",105,"# Abstract\n# Introduction\n## Partial monitoring and regret\n## Stochastic and contextual settings\n## Contributions","[{\"question\":\"What does the partial monitoring (PM) framework assume about feedback in each round?\",\"answer\":\"The agent selects an action while the environment simultaneously selects an outcome, but the agent does not observe the loss directly. Instead, it receives a feedback signal that is only partially informative about the unobserved outcome.\"},{\"question\":\"How does this paper extend regret-guarantee strategies for stochastic partial monitoring?\",\"answer\":\"It introduces a new class of strategies that randomize deterministic confidence bounds, enabling sub-linear regret guarantees in settings where existing stochastic strategies are not applicable.\"},{\"question\":\"What problem application is proposed to encourage adoption of the PM framework?\",\"answer\":\"The paper designs a real-world use case for monitoring the error rate of a deployed classification system using the PM formulation.\"}]",1784195464,91,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"randomized-confidence-bounds-for-stochastic-partial-monitoring","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/randomized-confidence-bounds-for-stochastic-partial-monitoring/84411/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What does the partial monitoring (PM) framework assume about feedback in each round?","Question",{"text":75,"@type":76},"The agent selects an action while the environment simultaneously selects an outcome, but the agent does not observe the loss directly. Instead, it receives a feedback signal that is only partially informative about the unobserved outcome.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does this paper extend regret-guarantee strategies for stochastic partial monitoring?",{"text":80,"@type":76},"It introduces a new class of strategies that randomize deterministic confidence bounds, enabling sub-linear regret guarantees in settings where existing stochastic strategies are not applicable.",{"name":82,"@type":73,"acceptedAnswer":83},"What problem application is proposed to encourage adoption of the PM framework?",{"text":84,"@type":76},"The paper designs a real-world use case for monitoring the error rate of a deployed classification system using the PM formulation.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]