[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"detail-sidebar-cat-0-en-105":3,"doc-seo-128845-105":59,"doc-detail-128845-en":130},{"code":4,"msg":5,"data":6},0,"success",[7,13,18,23,28,33,38,43,48,51,55],{"id":8,"doc_module":4,"doc_module_name":9,"category_name":10,"show_sort_weight":11,"slug":12},1,"Document","Story & Novel",90,"story-novel",{"id":14,"doc_module":4,"doc_module_name":9,"category_name":15,"show_sort_weight":16,"slug":17},2,"Literature",80,"literature",{"id":19,"doc_module":4,"doc_module_name":9,"category_name":20,"show_sort_weight":21,"slug":22},4,"Exam",70,"exam",{"id":24,"doc_module":4,"doc_module_name":9,"category_name":25,"show_sort_weight":26,"slug":27},5,"Comic",60,"comic",{"id":29,"doc_module":4,"doc_module_name":9,"category_name":30,"show_sort_weight":31,"slug":32},6,"Technology",50,"technology",{"id":34,"doc_module":4,"doc_module_name":9,"category_name":35,"show_sort_weight":36,"slug":37},7,"Healthcare",40,"healthcare",{"id":39,"doc_module":4,"doc_module_name":9,"category_name":40,"show_sort_weight":41,"slug":42},8,"Research & Report",30,"research-report",{"id":44,"doc_module":4,"doc_module_name":9,"category_name":45,"show_sort_weight":46,"slug":47},9,"Religion & Spirituality",20,"religion-spirituality",{"id":46,"doc_module":4,"doc_module_name":9,"category_name":49,"show_sort_weight":46,"slug":50},"World Cup","world-cup",{"id":52,"doc_module":4,"doc_module_name":9,"category_name":53,"show_sort_weight":52,"slug":54},10,"Lifestyle","lifestyle",{"id":56,"doc_module":4,"doc_module_name":9,"category_name":57,"show_sort_weight":24,"slug":58},19,"General","general",{"code":4,"msg":60,"data":61},"ok",{"site_id":62,"language":63,"slug":64,"title":65,"keywords":66,"description":67,"schema_data":68,"social_meta":123,"head_meta":125,"extra_data":127,"updated_unix":129},105,"en","hierarchical-reinforcement-learning-with-targeted-causal-interventions","Hierarchical Reinforcement Learning with Targeted Causal Interventions","","Hierarchical reinforcement learning addresses long-horizon tasks with sparse rewards by decomposing problems into subgoals, but it remains difficult to discover effective subgoal structure and use it to reach the final objective. This work models subgoal relationships as a causal graph, learns the causal structure via a causal discovery algorithm, and uses it to prioritize which subgoals to intervene on during exploration. Targeted interventions reduce training cost by improving policy efficiency. The paper also provides formal theoretical analysis for tree structures and a variant of Erdős–Rényi random graphs, supported by experiments showing lower training cost than prior approaches.",{"@graph":69,"@context":122},[70,84,105],{"@type":71,"itemListElement":72},"BreadcrumbList",[73,77,79,82],{"item":74,"name":75,"@type":76,"position":8},"https://docshare.wps.com","Home","ListItem",{"item":78,"name":9,"@type":76,"position":14},"https://docshare.wps.com/document/",{"item":80,"name":40,"@type":76,"position":81},"https://docshare.wps.com/document/research-report/",3,{"item":83,"name":65,"@type":76,"position":19},"https://docshare.wps.com/document/hierarchical-reinforcement-learning-with-targeted-causal-interventions/128845/",{"url":83,"name":65,"@type":85,"image":86,"author":91,"headline":65,"publisher":94,"fileFormat":97,"inLanguage":63,"description":67,"dateModified":98,"datePublished":99,"encodingFormat":97,"isAccessibleForFree":100,"interactionStatistic":101},"DigitalDocument",{"url":87,"@type":88,"width":89,"height":90},"https://docshare.wps.com/thumbnails/hierarchical-reinforcement-learning-with-targeted-causal-interventions/128845.png","ImageObject",300,407,{"name":92,"@type":93},"Aria","Person",{"url":74,"name":95,"@type":96},"DocShare","Organization","application/pdf","2026-09-18","2026-08-06",true,{"@type":102,"interactionType":103,"userInteractionCount":24},"InteractionCounter",{"@type":104},"ViewAction",{"@type":106,"mainEntity":107},"FAQPage",[108,114,118],{"name":109,"@type":110,"acceptedAnswer":111},"What problem does the paper address in hierarchical reinforcement learning?","Question",{"text":112,"@type":113},"It targets two core issues: efficiently discovering the hierarchy among subgoals and leveraging that structure to achieve the final goal in long-horizon, sparse-reward tasks.","Answer",{"name":115,"@type":110,"acceptedAnswer":116},"How does the proposed method use causal graphs during exploration?",{"text":117,"@type":113},"It learns subgoal structure as a causal graph and uses the learned causal model to prioritize which subgoals to intervene on, rather than selecting interventions at random.",{"name":119,"@type":110,"acceptedAnswer":120},"What theoretical results are provided, and how do experiments validate them?",{"text":121,"@type":113},"The paper offers formal analysis for tree structures and for a variant of Erdős–Rényi random graphs. Experiments on HRL tasks further show the framework outperforms existing methods in training cost.","https://schema.org",{"og:url":83,"og:type":124,"og:title":65,"og:site_name":95,"og:description":67},"article",{"robots":126,"canonical":83},"index,follow",{"doc_id":128,"site_id":62},128845,1786003842,{"code":4,"msg":5,"data":131},{"doc_id":128,"user_id":132,"nickname":92,"user_avatar":133,"doc_module":4,"category_id":39,"category_name":40,"doc_title":65,"doc_description":67,"doc_content":134,"file_id":135,"file_url":136,"file_type":137,"file_size":138,"view_count":24,"is_deleted":4,"is_public":8,"is_downloadable":8,"audit_status":8,"page_count":139,"language":140,"language_code":63,"site_id":62,"html_lang":63,"table_of_contents":141,"faqs":142,"seo_title":143,"seo_description":67,"update_tm":129,"read_time":144},2336474459895,"https://ap-avatar.wpscdn.com/avatar/22000baeef7a5ed0655?x-image-process=image/resize,m_fixed,w_180,h_180&k=1786071322749376916","Hierarchical Reinforcement Learning with Targeted Causal Interventions  \nSadegh Khorasani 1 Saber Salehkaleybar 2 Negar Kiyavash 3 Matthias Grossglauser 1  \nAbstract  \nHierarchical reinforcement learning (HRL) improves the efficiency of long-horizon reinforcement-learning tasks with sparse rewards by decomposing the task into a hierarchy of subgoals. The main challenge of HRL is efficient discovery of the hierarchical structure among subgoals and utilizing this structure to achieve the final goal. We address this challenge by modeling the subgoal structure as a causal graph and propose a causal discovery algorithm to learn it. Additionally, rather than intervening on the subgoals at random during exploration, we harness the discovered causal model to prioritize subgoal interventions based on their importance in attaining the final goal. These targeted interventions result in a significantly more efficient policy in terms of the training cost.  \nUnlike previous work on causal HRL, which lacked theoretical analysis, we provide a formal analysis of the problem. Specifically, for tree structures and, for a variant of Erds-Rnyi random graphs, our approach results in remarkable improvements. Our experimental results on HRL tasks also illustrate that our proposed framework outperforms existing work in terms of training cost.  \n1. Introduction  \nIn traditional reinforcement learning (RL), an agent is typically required to solve a specific task based on immediate feedback from  the environment (Sutton, 2018) . However, in  \n1 School of Computer and Communication Sciences, EPFL, Lausanne, Switzerland 2Leiden Institute of Advanced Computer Science (LIACS), Leiden University, Leiden, The Netherlands 3 College of Management of Technology, EPFL, Lausanne, Switzerland. Correspondence to: Sadegh Khorasani \u003C[sadegh.khorasani@epfl.ch](sadegh.khorasani@epfl.ch) >, Saber Salehkaleybar \u003C[s.salehkaleybar@liacs.leidenuniv.nl](s.salehkaleybar@liacs.leidenuniv.nl) >, Negar Kiyavash \u003C[negar.kiyavash@epfl.ch](negar.kiyavash@epfl.ch) >, Matthias Grossglauser \u003C[matthias.grossglauser@epfl.ch](matthias.grossglauser@epfl.ch) >.  \nProceedings of the 42 nd International Conference on Machine Learning, Vancouver, Canada. PMLR 267, 2025 . Copyright 2025 by the author(s) .  \nmany real-world applications, the agent faces tasks where rewards are sparse and delayed. This poses a significant challenge to traditional RL since the agent must take many actions without receiving an immediate reward. For instance, in maze-solving tasks, the agent might only receive a reward once it finds a path to a specific exit or a location within the maze. Hierarchical reinforcement learning (HRL) allows the agent to break down the problem into subtasks, which is helpful in environments where achieving goals requires a long horizon (Bacon et al., 2017 ; Barto & Mahadevan, 2003 ; Eysenbach et al., 2019 ; Le et al., 2018) .  \nHierarchical policies are widely used in HRL to manage the complexity of long-horizon and sparse-reward tasks. However, learning hierarchical policies introduces significant challenges. To address these issues, many methods are proposed to improve learning efficiency (Kulkarni et al., 2016 ; Levy et al., 2017 ; Vezhnevets et al., 2017) . The high-level policy does not deal with primitive actions but instead sets subgoals or options (Sutton et al., 1999) that the lower-level policies must achieve. Lower-level policies are trained to achieve subgoals assigned by the high-level policy. In order to discover meaningful subgoals and their relationships, it is essential to ensure that both high-level and low-level policies explore the environment efficiently. However, naive exploration of the entire goal space is not sample efficient in high-dimensional goal spaces. Some studies in the literature aim to improve sample efficiency and scalability by employing off-policy techniques (Nachum et al., 2018b), utilizing representation learning (Nachum et al., 2018a), or restricti","cbCaihYNsM11L3Gm","https://ap.wps.com/l/cbCaihYNsM11L3Gm","pdf",2389777,44,"English","# Abstract\n# Introduction\n## Sparse and delayed rewards in reinforcement learning\n## Challenges in discovering hierarchical structure\n## Causal HRL and limitations of prior work\n## Proposed HRC framework and contributions","[{\"question\":\"What problem does the paper address in hierarchical reinforcement learning?\",\"answer\":\"It targets two core issues: efficiently discovering the hierarchy among subgoals and leveraging that structure to achieve the final goal in long-horizon, sparse-reward tasks.\"},{\"question\":\"How does the proposed method use causal graphs during exploration?\",\"answer\":\"It learns subgoal structure as a causal graph and uses the learned causal model to prioritize which subgoals to intervene on, rather than selecting interventions at random.\"},{\"question\":\"What theoretical results are provided, and how do experiments validate them?\",\"answer\":\"The paper offers formal analysis for tree structures and for a variant of Erdős–Rényi random graphs. Experiments on HRL tasks further show the framework outperforms existing methods in training cost.\"}]","Hierarchical Reinforcement Learning with Targeted Causal Interventions | PDF",111]