[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81751-en":3,"doc-seo-81751-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81751,4398048949847,"Eliana","https://ap-avatar.wpscdn.com/avatar/400002536579ef2da7f?_k=1778318612642679267",8,"Research & Report","A Contextual-Bandit Oversight Game with Two-Sided Informational Asymmetry","Runtime human oversight of an AI agent is studied under two-direction private information: the human privately knows her reward function while the AI privately knows the quality of the action it proposes. The work builds on Cooperative Inverse Reinforcement Learning and the Oversight Game, introducing a contextual-bandit team game with a play/ask/trust/oversee interface. A two-sided slab of avoidable harm quantifies cases where AI knows harm yet a myopic, trusting human does not oversee. Non-credible oversight communication explains the gap, analyzed across repeated rounds via passive learning and active signaling.","arXiv :2607 .00 155v 1 [ cs .AI] 30 Jun 2026  \nA Contextual-Bandit Oversight Game with Two-Sided Informational Asymmetry  \nYunjin Tong  \nStanford Graduate School of Business  \nAbstract  \nWe study runtime human oversight of an AI agent when private information runs in both directions: the human privately knows her reward function, while the AI privately knows the quality of the action it proposes. This is the kind of asymmetry that arises naturally when an autonomous robot or software agent has inspected a situation its human supervisor cannot directly assess. Building on Cooperative Inverse Reinforcement Learning (CIRL) and the Oversight Game, we introduce a contextual-bandit team game with two-sided asymmetric information and a play/ask/trust/oversee interface. The bandit structure removes physical state transitions and thereby yields exact one-shot characterizations that would remain conjectural in the full POMDP setting, though the common belief remains a dynamically controlled state across rounds. We give two one-shot characterizations, a team optimum and a behaviorally natural myopic rule, whose gap is a “slab” of avoidable harm: a region in which the AI privately knows the proposed action is harmful and shutdown would help, yet a myopic human, trusting her prior, declines to oversee. We show this gap is the price of non-credible oversight communication, and give a partial analysis of how it resolves dynamically over repeated rounds through passive learning and active signaling with a one-period-lagged oversight response.  \n1 Introduction  \nA central problem in deploying autonomous agents, robotic or software, is calibrating when a human supervisor should intervene. As such agents take on consequential tasks, from a warehouse robot grasping a loaded shelf to a coding agent refactoring production software, the question of when a human should step in and override becomes a design problem in its own right: intervene too rarely and harmful actions slip through, intervene too often and the agent’s autonomy is wasted on costly and unnecessary oversight.  \nTwo lines of prior work frame the building blocks we combine. Cooperative Inverse Reinforcement Learning (CIRL) [1] casts human–AI interaction as a shared-reward game in which the AI is uncertain about the human’s preferences and must learn them through interaction. CIRL integrates preference learning with action selection and can generate active learning, active teaching, and communicative behavior; the human’s private reward parameter is the hidden information, anda common posterior over that parameter is the sufficient statistic for optimal play. What CIRL does not explicitly model is the runtime play/ask/trust/oversee interface studied here, nor an AI-private proposal-quality parameter that the human cannot observe. Its uncertainty is one-sided: it models“what does the human want?” but never “what does the AI know about the world that the human does not?” The Off-Switch Game [2] introduces runtime deferral as an explicit object of study, but only in a single-shot setting. The Oversight Game [3] supplies a runtime interface of the kind we use, in which an AI proposes an action, a human may override, and interaction costs make the  \ndecision nontrivial, but it is a Markov game under full information, with neither preference nor model uncertainty.  \nThis paper develops a model in which private information runs in both directions, and in which deferral is a runtime decision. The motivating observation is that an embodied or autonomous agent routinely knows things about the consequences of its own proposed actions that its supervisor cannot directly observe: a robot that has physically inspected its workspace, or a software agent that has read a codebase, has private knowledge of failure modes the human cannot see. This asymmetry runs opposite to CIRL’s. We therefore study a setting with two-sided private information, in which the human privately knows her reward type θ and the","cbCaiotDOZMygTGH","https://ap.wps.com/l/cbCaiotDOZMygTGH","pdf",414189,3,1,17,"English","en",105,"# Introduction\n## Background and Motivation\n## Model and Two-Sided Private Information\n## Methodological Choice: Contextual-Bandit Reduction\n## Contributions","[{\"question\":\"What two-sided private information does the paper assume in the oversight setting?\",\"answer\":\"The human privately knows her reward function, while the AI privately knows the quality (under an observation-model type) of the action it proposes.\"},{\"question\":\"How does the contextual-bandit model simplify the otherwise difficult full Markov/POMDP treatment?\",\"answer\":\"It removes physical state transitions, making the one-shot correction value simpler and enabling exact one-shot characterizations, while keeping the common belief as a dynamically controlled state across rounds.\"},{\"question\":\"What does the paper call the “slab” of avoidable harm, and why does it matter?\",\"answer\":\"It is a region where the AI privately knows the proposed action is harmful and shutdown would help, yet a myopic human who trusts her prior chooses not to oversee, illustrating costs from non-credible oversight communication.\"}]",1784175824,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"a-contextual-bandit-oversight-game-with-two-sided-informational-asymmetry","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/a-contextual-bandit-oversight-game-with-two-sided-informational-asymmetry/81751/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What two-sided private information does the paper assume in the oversight setting?","Question",{"text":75,"@type":76},"The human privately knows her reward function, while the AI privately knows the quality (under an observation-model type) of the action it proposes.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the contextual-bandit model simplify the otherwise difficult full Markov/POMDP treatment?",{"text":80,"@type":76},"It removes physical state transitions, making the one-shot correction value simpler and enabling exact one-shot characterizations, while keeping the common belief as a dynamically controlled state across rounds.",{"name":82,"@type":73,"acceptedAnswer":83},"What does the paper call the “slab” of avoidable harm, and why does it matter?",{"text":84,"@type":76},"It is a region where the AI privately knows the proposed action is harmful and shutdown would help, yet a myopic human who trusts her prior chooses not to oversee, illustrating costs from non-credible oversight communication.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]