[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"detail-sidebar-cat-0-en-105":3,"doc-seo-137567-105":59,"doc-detail-137567-en":130},{"code":4,"msg":5,"data":6},0,"success",[7,13,18,23,28,33,38,43,48,51,55],{"id":8,"doc_module":4,"doc_module_name":9,"category_name":10,"show_sort_weight":11,"slug":12},1,"Document","Story & Novel",90,"story-novel",{"id":14,"doc_module":4,"doc_module_name":9,"category_name":15,"show_sort_weight":16,"slug":17},2,"Literature",80,"literature",{"id":19,"doc_module":4,"doc_module_name":9,"category_name":20,"show_sort_weight":21,"slug":22},4,"Exam",70,"exam",{"id":24,"doc_module":4,"doc_module_name":9,"category_name":25,"show_sort_weight":26,"slug":27},5,"Comic",60,"comic",{"id":29,"doc_module":4,"doc_module_name":9,"category_name":30,"show_sort_weight":31,"slug":32},6,"Technology",50,"technology",{"id":34,"doc_module":4,"doc_module_name":9,"category_name":35,"show_sort_weight":36,"slug":37},7,"Healthcare",40,"healthcare",{"id":39,"doc_module":4,"doc_module_name":9,"category_name":40,"show_sort_weight":41,"slug":42},8,"Research & Report",30,"research-report",{"id":44,"doc_module":4,"doc_module_name":9,"category_name":45,"show_sort_weight":46,"slug":47},9,"Religion & Spirituality",20,"religion-spirituality",{"id":46,"doc_module":4,"doc_module_name":9,"category_name":49,"show_sort_weight":46,"slug":50},"World Cup","world-cup",{"id":52,"doc_module":4,"doc_module_name":9,"category_name":53,"show_sort_weight":52,"slug":54},10,"Lifestyle","lifestyle",{"id":56,"doc_module":4,"doc_module_name":9,"category_name":57,"show_sort_weight":24,"slug":58},19,"General","general",{"code":4,"msg":60,"data":61},"ok",{"site_id":62,"language":63,"slug":64,"title":65,"keywords":66,"description":67,"schema_data":68,"social_meta":123,"head_meta":125,"extra_data":127,"updated_unix":129},105,"en","interactive-learning-from-policy-dependent-human-feedback","Interactive Learning from Policy-Dependent Human Feedback","","This paper investigates interactively learning behaviors conveyed by a human teacher through positive and negative feedback. Prior research typically assumes feedback depends on the behavior being taught but is independent of the learner’s current policy. Empirical results show this assumption is wrong: trainers’ feedback (positive or negative) is influenced by the learner’s policy. The paper proposes Convergent Actor-Critic by Humans (COACH), which learns from policy-dependent feedback and converges to a local optimum, validated on a physical robot.",{"@graph":69,"@context":122},[70,84,105],{"@type":71,"itemListElement":72},"BreadcrumbList",[73,77,79,82],{"item":74,"name":75,"@type":76,"position":8},"https://docshare.wps.com","Home","ListItem",{"item":78,"name":9,"@type":76,"position":14},"https://docshare.wps.com/document/",{"item":80,"name":40,"@type":76,"position":81},"https://docshare.wps.com/document/research-report/",3,{"item":83,"name":65,"@type":76,"position":19},"https://docshare.wps.com/document/interactive-learning-from-policy-dependent-human-feedback/137567/",{"url":83,"name":65,"@type":85,"image":86,"author":91,"headline":65,"publisher":94,"fileFormat":97,"inLanguage":63,"description":67,"dateModified":98,"datePublished":99,"encodingFormat":97,"isAccessibleForFree":100,"interactionStatistic":101},"DigitalDocument",{"url":87,"@type":88,"width":89,"height":90},"https://docshare.wps.com/thumbnails/interactive-learning-from-policy-dependent-human-feedback/137567.png","ImageObject",300,407,{"name":92,"@type":93},"Olivia Brown","Person",{"url":74,"name":95,"@type":96},"DocShare","Organization","application/pdf","2026-09-20","2026-08-22",true,{"@type":102,"interactionType":103,"userInteractionCount":52},"InteractionCounter",{"@type":104},"ViewAction",{"@type":106,"mainEntity":107},"FAQPage",[108,114,118],{"name":109,"@type":110,"acceptedAnswer":111},"What problem does the paper study in interactive human feedback learning?","Question",{"text":112,"@type":113},"It studies how an agent can learn behaviors communicated by a human teacher using positive and negative feedback during interaction.","Answer",{"name":115,"@type":110,"acceptedAnswer":116},"What key assumption in prior work does this paper challenge?",{"text":117,"@type":113},"It challenges the assumption that human feedback is independent of the learner’s current policy, showing that feedback changes with the learner’s policy.",{"name":119,"@type":110,"acceptedAnswer":120},"What is COACH and what does it achieve?",{"text":121,"@type":113},"COACH (Convergent Actor-Critic by Humans) is an algorithm for learning from policy-dependent human feedback, and it is shown to converge to a local optimum while successfully learning multiple behaviors on a physical robot.","https://schema.org",{"og:url":83,"og:type":124,"og:title":65,"og:site_name":95,"og:description":67},"article",{"robots":126,"canonical":83},"index,follow",{"doc_id":128,"site_id":62},137567,1787421224,{"code":4,"msg":5,"data":131},{"doc_id":128,"user_id":132,"nickname":92,"user_avatar":133,"doc_module":4,"category_id":39,"category_name":40,"doc_title":65,"doc_description":67,"doc_content":134,"file_id":135,"file_url":136,"file_type":137,"file_size":138,"view_count":52,"is_deleted":4,"is_public":8,"is_downloadable":8,"audit_status":8,"page_count":52,"language":139,"language_code":63,"site_id":62,"html_lang":63,"table_of_contents":140,"faqs":141,"seo_title":142,"seo_description":67,"update_tm":129,"read_time":143},16904993612988,"https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd","Interactive Learning from Policy-Dependent Human Feedback  \nJames MacGlashan 1 Mark K Ho 2 Robert Loftin 3 Bei Peng 4 Guan Wang 2 David L. Roberts 3  \nMatthew E. Taylor 4 Michael L. Littman 2  \nAbstract  \nThis paper investigates the problem of interactively learning behaviors communicated by a human teacher using positive and negative feedback. Much previous work on this problem has made the assumption that people provide feedback for decisions that is dependent on the behavior they are teaching and is independent from the learner's current policy. We present empirical results that show this assumption to be false—whether human trainers give a positive or negative feedback for a decision is inﬂuenced by the learner's current policy. Based on this insight, we introduce Convergent Actor-Critic by Humans (COACH), an algorithm for learning from policy-dependent feedback that converges to a local optimum. Finally, we demonstrate that COACH can successfully learn multiple behaviors on a physical robot.  \n1. Introduction  \nProgramming robots is very difﬁcult, in part because the real world is inherently rich and—to some degree—unpredictable. In addition, our expectations for physical agents are quite high and often difﬁcult to articulate. Nevertheless, for robots to have a signiﬁcant impact on the lives of individuals, even non-programmers need to be able to specify and customize behavior. Because of these complexities, relying on end-users to provide instructions to robots programmatically seems destined to fail.  \nReinforcement learning (RL) from human trainer feedback provides a compelling alternative to programming because agents can learn complex behavior from very simple positive and negative signals. Furthermore, real-world animal training is an existence proof that people can train complex  \n*Equal contribution 1Cogitai 2Brown University 3North Carolina State University 4Washington State University. Correspondence to: James MacGlashan \u003C[james@cogitai.com](james@cogitai.com) >.  \nProceedings of the 34 th International Conference on Machine Learning, Sydney, Australia, PMLR 70, 2017 . Copyright 2017 by the author(s) .  \nbehavior using these simple signals. Indeed, animals have been successfully trained to guide the blind, locate mines in the ocean, detect cancer or explosives, and even solve complex, multi-stage puzzles.  \nDespite success when learning from environmental reward, traditional reinforcement-learning algorithms have yielded limited success when the reward signal is provided by humans. This failure underscores the importance that algorithms for learning from humans are based on appropriate models of human-feedback. Indeed, much human-centered RL work has investigated and employed different models of human-feedback (Knox & Stone, 2009b ; Thomaz & Breazeal, 2006 ; 2007 ; 2008 ; Grifﬁth et al., 2013 ; Loftin et al., 2015) . Many of these algorithms leverage the observation that people tend to give feedback that is best interpreted as guidance on the policy the agent should be following, rather than as a numeric value to be maximized by the agent. However, these approaches assume models of feedback that are independent of the policy the agent is currently following. We present empirical results that demonstrate that this assumption is incorrect and further demonstrate cases in which policy-independent learning algorithms suffer from this assumption. Following this result, we present Convergent Actor-Critic by Humans (COACH), an algorithm for learning from policy-dependent human feedback. COACH is based on the insight that the advantage function (a value roughly corresponding to how much better or worse an action is compared to the current policy) provides a better model of human feedback, capturing human-feedback properties like diminishing returns, rewarding improvement, and giving 0-valued feedback a semantic meaning that combats forgetting. We compare COACH to other approaches in a simple domain with simulated feedba","cbCaigsx8BYgD4m8","https://ap.wps.com/l/cbCaigsx8BYgD4m8","pdf",410757,"English","# Introduction\n# Background","[{\"question\":\"What problem does the paper study in interactive human feedback learning?\",\"answer\":\"It studies how an agent can learn behaviors communicated by a human teacher using positive and negative feedback during interaction.\"},{\"question\":\"What key assumption in prior work does this paper challenge?\",\"answer\":\"It challenges the assumption that human feedback is independent of the learner’s current policy, showing that feedback changes with the learner’s policy.\"},{\"question\":\"What is COACH and what does it achieve?\",\"answer\":\"COACH (Convergent Actor-Critic by Humans) is an algorithm for learning from policy-dependent human feedback, and it is shown to converge to a local optimum while successfully learning multiple behaviors on a physical robot.\"}]","Interactive Learning from Policy-Dependent Human Feedback | PDF",25]