[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"detail-sidebar-cat-0-en-105":3,"doc-seo-128864-105":59,"doc-detail-128864-en":131},{"code":4,"msg":5,"data":6},0,"success",[7,13,18,23,28,33,38,43,48,51,55],{"id":8,"doc_module":4,"doc_module_name":9,"category_name":10,"show_sort_weight":11,"slug":12},1,"Document","Story & Novel",90,"story-novel",{"id":14,"doc_module":4,"doc_module_name":9,"category_name":15,"show_sort_weight":16,"slug":17},2,"Literature",80,"literature",{"id":19,"doc_module":4,"doc_module_name":9,"category_name":20,"show_sort_weight":21,"slug":22},4,"Exam",70,"exam",{"id":24,"doc_module":4,"doc_module_name":9,"category_name":25,"show_sort_weight":26,"slug":27},5,"Comic",60,"comic",{"id":29,"doc_module":4,"doc_module_name":9,"category_name":30,"show_sort_weight":31,"slug":32},6,"Technology",50,"technology",{"id":34,"doc_module":4,"doc_module_name":9,"category_name":35,"show_sort_weight":36,"slug":37},7,"Healthcare",40,"healthcare",{"id":39,"doc_module":4,"doc_module_name":9,"category_name":40,"show_sort_weight":41,"slug":42},8,"Research & Report",30,"research-report",{"id":44,"doc_module":4,"doc_module_name":9,"category_name":45,"show_sort_weight":46,"slug":47},9,"Religion & Spirituality",20,"religion-spirituality",{"id":46,"doc_module":4,"doc_module_name":9,"category_name":49,"show_sort_weight":46,"slug":50},"World Cup","world-cup",{"id":52,"doc_module":4,"doc_module_name":9,"category_name":53,"show_sort_weight":52,"slug":54},10,"Lifestyle","lifestyle",{"id":56,"doc_module":4,"doc_module_name":9,"category_name":57,"show_sort_weight":24,"slug":58},19,"General","general",{"code":4,"msg":60,"data":61},"ok",{"site_id":62,"language":63,"slug":64,"title":65,"keywords":66,"description":67,"schema_data":68,"social_meta":124,"head_meta":126,"extra_data":128,"updated_unix":130},105,"en","preference-elicitation-and-inverse-reinforcement-learning","Preference Elicitation and Inverse Reinforcement Learning","","We formulate inverse reinforcement learning as a preference elicitation problem and provide a principled Bayesian statistical framework. The approach generalizes prior Bayesian inverse reinforcement learning, producing a posterior over an agent’s preferences and policy, and optionally over the reward sequence, given observational data. We analyze how this Bayesian method relates to other inverse reinforcement learning techniques and validate it with experiments. Results show accurate preference recovery even when the observed policy is sub-optimal for its own preferences, yielding markedly improved policies compared with alternatives and demonstrated performance.",{"@graph":69,"@context":123},[70,84,106],{"@type":71,"itemListElement":72},"BreadcrumbList",[73,77,79,82],{"item":74,"name":75,"@type":76,"position":8},"https://docshare.wps.com","Home","ListItem",{"item":78,"name":9,"@type":76,"position":14},"https://docshare.wps.com/document/",{"item":80,"name":40,"@type":76,"position":81},"https://docshare.wps.com/document/research-report/",3,{"item":83,"name":65,"@type":76,"position":19},"https://docshare.wps.com/document/preference-elicitation-and-inverse-reinforcement-learning/128864/",{"url":83,"name":65,"@type":85,"image":86,"author":91,"headline":65,"publisher":94,"fileFormat":97,"inLanguage":63,"description":67,"dateModified":98,"datePublished":99,"encodingFormat":97,"isAccessibleForFree":100,"interactionStatistic":101},"DigitalDocument",{"url":87,"@type":88,"width":89,"height":90},"https://docshare.wps.com/thumbnails/preference-elicitation-and-inverse-reinforcement-learning/128864.png","ImageObject",300,407,{"name":92,"@type":93},"Noah","Person",{"url":74,"name":95,"@type":96},"DocShare","Organization","application/pdf","2026-09-19","2026-08-06",true,{"@type":102,"interactionType":103,"userInteractionCount":105},"InteractionCounter",{"@type":104},"ViewAction",11,{"@type":107,"mainEntity":108},"FAQPage",[109,115,119],{"name":110,"@type":111,"acceptedAnswer":112},"How does the paper connect inverse reinforcement learning with preference elicitation?","Question",{"text":113,"@type":114},"It defines inverse reinforcement learning as a preference elicitation problem, casting the task of inferring an agent’s preferences from observed behavior into a Bayesian statistical formulation.","Answer",{"name":116,"@type":111,"acceptedAnswer":117},"What does the Bayesian framework infer from observations?",{"text":118,"@type":114},"From observations of the agent interacting with a stochastic environment, it derives a posterior distribution over the agent’s preferences and policy, and optionally the obtained reward sequence.",{"name":120,"@type":111,"acceptedAnswer":121},"Can the method recover accurate preferences when the demonstrated policy is sub-optimal?",{"text":122,"@type":114},"Yes. The experiments indicate preferences can be determined accurately even if the observed policy is sub-optimal with respect to the agent’s own preferences, and the inferred policies improve performance substantially.","https://schema.org",{"og:url":83,"og:type":125,"og:title":65,"og:site_name":95,"og:description":67},"article",{"robots":127,"canonical":83},"index,follow",{"doc_id":129,"site_id":62},128864,1786004023,{"code":4,"msg":5,"data":132},{"doc_id":129,"user_id":133,"nickname":92,"user_avatar":134,"doc_module":4,"category_id":39,"category_name":40,"doc_title":65,"doc_description":67,"doc_content":135,"file_id":136,"file_url":137,"file_type":138,"file_size":139,"view_count":105,"is_deleted":4,"is_public":8,"is_downloadable":8,"audit_status":8,"page_count":140,"language":141,"language_code":63,"site_id":62,"html_lang":63,"table_of_contents":142,"faqs":143,"seo_title":144,"seo_description":67,"update_tm":130,"read_time":145},137451207643,"https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2","Preference Elicitation and Inverse Reinforcement Learning  \nConstantin A. Rothkopf1 and Christos Dimitrakakis2  \n1 Frankfurt Institute for Advanced Studies, Frankfurt, Germany  \n[rothkfopf@fias.uni-frankfurt.de](rothkfopf@fias.uni-frankfurt.de)  \n2 EPFL, Lausanne, Switzerland  \nchristos .dimitrakakis@epfl .ch  \nAbstract. We state the problem of inverse reinforcement learning in terms of preference elicitation, resulting in a principled (Bayesian) statistical formulation. This generalises previous work on Bayesian inverse reinforcement learning and allows us to obtain a posterior distribution on the agent’s preferences, policy and optionally, the obtained reward sequence, from observations. We examine the relation of the resulting approach to other statistical methods for inverse reinforcement learning via analysis and experimental results. We show that preferences can be determined accurately, even if the observed agent’s policy is sub-optimal with respect to its own preferences. In that case, signiﬁcantly improved policies with respect to the agent’s preferences are obtained, compared to both other methods and to the performance of the demonstrated policy.  \nKeywords: Inverse reinforcement learning, preference elicitation, decision theory, Bayesian inference.  \n1 Introduction  \nPreference elicitation is a well-known problem in statistical decision theory [10] . The goal is to determine, whether a given decision maker prefers some events to other events, and if so, by how much. The ﬁrst main assumption is that there exists a partial ordering among events, indicating relative preferences. Then the corresponding problem is to determine which events are preferred to which others. The second main assumption is the expected utility hypothesis. This posits that if we can assign a numerical utility to each event, such that events with larger utilities are preferred, then the decision maker’s preferred choice from a set of possible gambles will be the gamble with the highest expected utility. The corresponding problem is to determine the numerical utilities for a given decision maker.  \nPreference elicitation is also of relevance to cognitive science and behavioural psychology, e.g. for determining rewards implicit in behaviour [19] where a proper elicitation procedure may allow one to reach more robust experimental conclusions. There are also direct practical applications, such as user modelling for determining customer preferences [3] . Finally, by analysing the apparent preferences of an expert while performing a particular task, we may be able to discover  \nD. Gunopulos et al. (Eds.): ECML PKDD 2011, Part III, LNAI 6913, pp. 34–48, 2011 .  \n􀀂c Springer-Verlag Berlin Heidelberg 2011  \nPreference Elicitation and Inverse Reinforcement Learning 35  \nbehaviours that match or even surpass the performance of the expert [1] in the very same task.  \nThis paper uses the formal setting of preference elicitation to determine the preferences of an agent acting within a discrete-time stochastic environment. We assume that the agent obtains a sequence of (hidden to us) rewards from the environment and that its preferences have a functional form related to the rewards. We also suppose that the agent is acting nearly optimally (in a manner to be made more rigorous later) with respect to its preferences. Armed with this information, and observations from the agent’s interaction with the environment, we can determine the agent’s preferences and policy in a Bayesian framework. This allows us to generalise previous Bayesian approaches to inverse reinforcement learning.  \nIn order to do so, we deﬁne a structured prior on reward functions and policies. We then derive two diﬀerent Markov chain procedures for preference elicitation. The result of the inference is used to obtain policies that are signiﬁcantly improved with respect to the true preferences of the observed agent. We show that this can be achieved even with fairly generic sampling approaches. ","cbCaiohSnPcEbsh2","https://ap.wps.com/l/cbCaiohSnPcEbsh2","pdf",331114,15,"English","# Introduction\n## Preference elicitation in decision theory\n## Motivation and applications\n## Bayesian formulation and objectives\n# Formalisation of the Problem\n## Separating preferences from environment dynamics\n# Abstract statistical model\n# Joint estimation of preferences and policy\n# Related work\n# Comparative experiments","[{\"question\":\"How does the paper connect inverse reinforcement learning with preference elicitation?\",\"answer\":\"It defines inverse reinforcement learning as a preference elicitation problem, casting the task of inferring an agent’s preferences from observed behavior into a Bayesian statistical formulation.\"},{\"question\":\"What does the Bayesian framework infer from observations?\",\"answer\":\"From observations of the agent interacting with a stochastic environment, it derives a posterior distribution over the agent’s preferences and policy, and optionally the obtained reward sequence.\"},{\"question\":\"Can the method recover accurate preferences when the demonstrated policy is sub-optimal?\",\"answer\":\"Yes. The experiments indicate preferences can be determined accurately even if the observed policy is sub-optimal with respect to the agent’s own preferences, and the inferred policies improve performance substantially.\"}]","Preference Elicitation and Inverse Reinforcement Learning | PDF",38]