[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85041-en":3,"doc-seo-85041-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85041,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","Provably Optimal Learning Algorithms for Assistance Games","This paper studies an online variant of the assistance games framework for repeated cooperative interaction between an informed agent and an uninformed assistant over T timesteps. The human observes a latent world state, while the assistant observes only the human’s actions. The work introduces assistance regret and provides provably efficient decentralized learning algorithms achieving (1−1/e)-approximate assistance regret with near T3/4 scaling. It also proves hardness beyond the (1−1/e) factor and shows a pseudo-decentralized shared-random-string extension reaching near-optimal T1/2 rates.","arXiv :2607 .080 12v 1 [ cs .LG] 9 Jul 2026  \nProvably Optimal Learning Algorithms for Assistance Games  \nNivasini Ananthakrishnan, Mark Bedaywi, Michael I. Jordan, Stuart Russell, and Nika  \nHaghtalab  \nUniversity of California, Berkeley  \n{nivasini,mark_bedaywi,michael_jordan,russell,[nika}@berkeley. edu](nika}@berkeley. edu)  \nAbstract  \nThis paper studies an online variant of the assistance games framework, where an informed agent and an uninformed agent repeatedly interact over T timesteps to optimize a common reward function. While the informed agent (the human) observes a latent state of the world, the uninformed agent (the assistant) observes only the human’s actions. We provide the first provably efficient learning algorithms for repeated assistance games. We introduce the notion of assistance regret: the gap between the cumulative utility of interactions and that of the optimal joint policies in hindsight, which map latent states to action pairs. We present decentralized algorithms for both the human and the assistant that achieve a (1 − 1/e)-approximate assistance regret rate of ˜O 􀀐T3/4􀀑 , with runtime polynomial in the size of the action and state spaces. These algorithms are general; in particular, they accommodate any no-regret algorithm for the assistant. We prove that achieving a regret approximation factor better than (1 − 1/e) is computationally intractable. Furthermore, we demonstrate how these generic no-regret algorithms can be tailored to a pseudo-decentralized setting—using a shared random string—to achieve a rate of ˜O 􀀐T 1/2􀀑 , optimal up to logarithmic factors.  \n1 Introduction  \nConsider a repeated interaction between two cooperative agents who share a common objective but have asymmetric access to information. One agent observes a changing latent state—such as a preference, type, or objective—while the other must act based only on indirect signals produced by the first agent. Such settings arise naturally in abstractions of human–assistant interaction, cooperative multi-agent systems, training of assistive AI, and emergent communication, where the former agent represents a human principal with private preferences and the latter represents an AI assistant seeking to act on the human’s behalf.  \nDespite the absence of strategic misalignment in utilities, coordination in these environments poses a fundamental challenge. On the one hand, the lack of a shared frame of reference or common language means that the agents must learn how to communicate by observing each other’s actions over time. On the other hand, actions simultaneously generate utility, causing them to serve a dual role as instruments for achieving reward and signals for conveying the latent state. This tension between informativeness and utility lies at the core of the problem.  \nOne theoretical perspective for capturing these challenges is the assistance games framework, also known as Cooperative Inverse Reinforcement Learning [Hadfield-Menell et al. , 2016] . This framework models interactions between a human and an assistive AI system as a cooperative game with partial observability. Prior work has studied equilibria and structural properties of optimal assistive policies—for example, tradeoffs between the informativeness and utility of actions—but the algorithmic problem of computing and learning such policies has been unexplored.  \nIn this paper, we introduce an online variant of assistance games and give computationally efficient, decentralized algorithms for learning near-optimal policies for both the human and the assistant.  \nOur model captures repeated interaction between an informed agent (the human) and an uninformed agent (the assistant) . In each round t, a latent state θ (t) ∈ Θ—possibly drawn from a nonstationary process—is realized and observed only by the human. This state represents private information about the human’spreferences or objectives. The human selects an action aH(t) from a finite action space AH of size MH ","cbCaip0C4VWkNP4t","https://ap.wps.com/l/cbCaip0C4VWkNP4t","pdf",810612,1,31,"English","en",105,"# Abstract\n# Introduction\n## Problem setup: informed human and uninformed assistant\n## Assistance regret and performance measure\n## Main contributions and theoretical results\n## Technical approach: reduction and regret decomposition","[{\"question\":\"What defines assistance games in this paper’s online setting?\",\"answer\":\"Two cooperative agents interact for T timesteps with partial observability: the informed human observes a latent state, while the assistant only observes the human’s actions and then chooses its own action.\"},{\"question\":\"How is performance measured?\",\"answer\":\"The paper uses assistance regret, defined as the gap between the agents’ cumulative reward and the reward of the best joint policies in hindsight mapping latent states to action pairs.\"},{\"question\":\"What are the main algorithmic and theoretical results?\",\"answer\":\"It provides first provably efficient decentralized learning algorithms achieving (1−1/e)-approximate assistance regret with near O(T^{3/4}) rate, plus a pseudo-decentralized shared-random-string variant achieving near-optimal O(T^{1/2}) up to logarithmic factors; it also shows achieving an approximation factor better than (1−1/e) is computationally intractable.\"}]",1784200565,78,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"provably-optimal-learning-algorithms-for-assistance-games","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/provably-optimal-learning-algorithms-for-assistance-games/85041/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What defines assistance games in this paper’s online setting?","Question",{"text":75,"@type":76},"Two cooperative agents interact for T timesteps with partial observability: the informed human observes a latent state, while the assistant only observes the human’s actions and then chooses its own action.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is performance measured?",{"text":80,"@type":76},"The paper uses assistance regret, defined as the gap between the agents’ cumulative reward and the reward of the best joint policies in hindsight mapping latent states to action pairs.",{"name":82,"@type":73,"acceptedAnswer":83},"What are the main algorithmic and theoretical results?",{"text":84,"@type":76},"It provides first provably efficient decentralized learning algorithms achieving (1−1/e)-approximate assistance regret with near O(T^{3/4}) rate, plus a pseudo-decentralized shared-random-string variant achieving near-optimal O(T^{1/2}) up to logarithmic factors; it also shows achieving an approximation factor better than (1−1/e) is computationally intractable.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]