[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122920-en":3,"doc-seo-122920-105":30,"detail-sidebar-cat-0-en-105":95},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122920,137441390410,"Hazel","https://ap-avatar.wpscdn.com/avatar/2000252f4ab5702993?_k=1776741390130283984",8,"Research & Report","Offline Inverse RL - New Solution Concepts and Provably Efficient Algorithms","Inverse reinforcement learning (IRL) seeks to recover an expert’s reward function from demonstrated behavior, but the task is inherently ill-posed because many rewards can explain the same trajectories. To address this, IRL is reframed as estimating the feasible reward set rather than choosing a single reward. Existing work focuses mostly on online learning where the agent can query and explore, which is unrealistic for practical offline applications. This paper introduces a feasible reward set tailored to offline data coverage limits, analyzes estimation complexity, and proposes efficient algorithms IRLO and PIRLO, with pessimism ensuring inclusion monotonicity.","Offline Inverse RL:  \nNew Solution Concepts and Provably Efficient Algorithms  \nFilippo Lazzati 1 Mirco Mutti 2 Alberto Maria Metelli 1  \nAbstract  \nInverse reinforcement learning (IRL) aims to recover the reward function of an expert agent from demonstrations of behavior. It is well-known that the IRL problem is fundamentally ill-posed, i.e., many reward functions can explain the demonstrations. For this reason, IRL has been recently reframed in terms of estimating the feasible reward set (Metelli et al., 2021), thus, postponing the selection of a single reward. However, so far, the available formulations and algorithmic solutions have been proposed and analyzed mainly for the online setting, where the learner can interact with the environment and query the expert at will. This is clearly unrealistic in most practical applications, where the availability of an offline dataset is a much more common scenario. In this paper, we introduce a novel notion of feasible reward set capturing the opportunities and limitations of the offline setting and we analyze the complexity of its estimation. This requires the introduction of an original learning framework that copes with the intrinsic difficulty of the setting, for which the data coverage is not under control. Then, we propose two computationally and statistically efficient algorithms, IRLO and PIRLO, for addressing the problem. In particular, the latter adopts a specific form of pessimism to enforce the novel, desirable property of inclusion monotonicity of the delivered feasible set. With this work, we aim to provide a panorama of the challenges of the offline IRL problem and how they can be fruitfully addressed.  \n1Politecnico di Milano, Milan, Italy 2Technion, Haifa, Israel. Correspondence to: Filippo Lazzati \u003C[filippo.lazzati@polimi.it](filippo.lazzati@polimi.it) >.  \nProceedings of the 41 st International Conference on Machine Learning, Vienna, Austria. PMLR 235, 2024 . Copyright 2024 by the author(s) .  \n1. Introduction  \nInverse reinforcement learning (IRL), also called inverse optimal control, consists of recovering a reward function from expert’s demonstrations (Russell, 1998) . Specifically, the reward is required to be compatible with the expert’s behavior, i.e., it shall make the expert’s policy optimal. As pointed out in Arora & Doshi (2018), IRL allows mitigating the challenging task of the manual specification of the reward function, thanks to the presence of demonstrations, and provides an effective method for imitation learning (Osa et al., 2018) . In opposition to mere behavioral cloning, IRL allows focusing on the expert intent (instead of behavior), and, for this reason, it has the potential to reveal the underlying objectives that drive the expert’s choices. In this sense, IRL enables interpretability, improving the interaction with the expert by explaining and predicting its behavior, and transferability, as the reward (more than a policy) can be employed under environment shifts (Adams et al., 2022) .  \nOne of the main concerns of IRL is that the problem is inherently ill-posed or ambiguous (Ng & Russell, 2000), i.e., there exists a variety of reward functions compatible with expert’s demonstrations. In the literature, many criteria for the selection of a single reward among the compatible ones were proposed (e.g., Ng & Russell, 2000 ; Ratliff et al., 2006 ; Ziebart et al., 2008 ; Boularias et al., 2011) . Nevertheless, the ambiguity issue has limited the theoretical understanding of the IRL problem for a long time.  \nRecently, IRL has been reframed by Metelli et al. (2021) into the problem of computing the set of all rewards compatible with expert’s demonstrations, named feasible reward set (or just feasible set) . By postponing the choice of a specific reward within the feasible set, this formulation has opened the doors to a new perspective that has enabled a deeper theoretical understanding of the IRL problem. The majority of previous works on the recons","cbCaihQVGDK9BAOK","https://ap.wps.com/l/cbCaihQVGDK9BAOK","pdf",791923,1,67,"English","en",105,"# 1. Introduction\n# Offline setting and feasible reward sets\n# Proposed algorithms: IRLO and PIRLO","[{\"question\":\"What problem does inverse reinforcement learning (IRL) address in this work?\",\"answer\":\"IRL aims to recover an expert’s reward function from demonstrations, ensuring the recovered reward is compatible with the expert’s behavior.\"},{\"question\":\"Why is IRL considered ill-posed, and what reformulation is used here?\",\"answer\":\"Multiple reward functions can explain the same demonstrations, creating ambiguity. The paper uses the feasible reward set formulation to avoid selecting a single reward.\"},{\"question\":\"What changes when moving from online IRL to offline IRL?\",\"answer\":\"Offline learning relies on a fixed dataset of trajectories and cannot control exploration or query the expert, making data coverage limitations central to the learning target.\"},{\"question\":\"How do IRLO and PIRLO contribute to efficient offline feasible set estimation?\",\"answer\":\"The paper proposes two computationally and statistically efficient algorithms, where PIRLO uses a pessimistic design to enforce inclusion monotonicity of the delivered feasible set.\"}]","Offline Inverse RL - New Solution Concepts and Provably Efficient Algorithms | PDF",1785813671,169,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":90,"head_meta":92,"extra_data":94,"updated_unix":28},"offline-inverse-rl-new-solution-concepts-and-provably-efficient-algorithms","",{"@graph":36,"@context":89},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/offline-inverse-rl-new-solution-concepts-and-provably-efficient-algorithms/122920/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81,85],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does inverse reinforcement learning (IRL) address in this work?","Question",{"text":75,"@type":76},"IRL aims to recover an expert’s reward function from demonstrations, ensuring the recovered reward is compatible with the expert’s behavior.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why is IRL considered ill-posed, and what reformulation is used here?",{"text":80,"@type":76},"Multiple reward functions can explain the same demonstrations, creating ambiguity. The paper uses the feasible reward set formulation to avoid selecting a single reward.",{"name":82,"@type":73,"acceptedAnswer":83},"What changes when moving from online IRL to offline IRL?",{"text":84,"@type":76},"Offline learning relies on a fixed dataset of trajectories and cannot control exploration or query the expert, making data coverage limitations central to the learning target.",{"name":86,"@type":73,"acceptedAnswer":87},"How do IRLO and PIRLO contribute to efficient offline feasible set estimation?",{"text":88,"@type":76},"The paper proposes two computationally and statistically efficient algorithms, where PIRLO uses a pessimistic design to enforce inclusion monotonicity of the delivered feasible set.","https://schema.org",{"og:url":52,"og:type":91,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":93,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":96},[97,101,105,109,114,119,124,127,132,135,139],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},"Exam",70,"exam",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},5,"Comic",60,"comic",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},6,"Technology",50,"technology",{"id":120,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":122,"slug":123},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":125,"slug":126},30,"research-report",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":130,"slug":131},9,"Religion & Spirituality",20,"religion-spirituality",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":130,"slug":134},"World Cup","world-cup",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":136,"slug":138},10,"Lifestyle","lifestyle",{"id":140,"doc_module":4,"doc_module_name":46,"category_name":141,"show_sort_weight":110,"slug":142},19,"General","general"]