[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122371-en":3,"doc-seo-122371-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122371,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",6,"Technology","Hierarchies of Reward Machines - curriculum-based hierarchical reward machines for reinforcement learning","Reward machines (RMs) provide a finite-state machine formalism to represent reinforcement learning reward functions, where edges encode task subgoals via high-level events and associated rewards. Their structure supports decomposing long-horizon and sparse-reward tasks into simpler, independently solvable subtasks. This work introduces a hierarchy of RMs (HRMs) by enabling an RM to call other RMs, composing complex behavior. Experiments show handcrafted HRMs accelerate convergence versus flat HRMs, and learning HRMs is practical even when equivalent flat representations are infeasible.","Hierarchies of Reward Machines  \nDaniel Furelos-Blanco 1 Mark Law 1 2 Anders Jonsson 3 Krysia Broda 1 Alessandra Russo 1  \nAbstract  \nReward machines (RMs) are a recent formalism for representing the reward function of a reinforcement learning task through a finite-state machine whose edges encode subgoals of the task using high-level events. The structure of RMs enables the decomposition of a task into simpler and independently solvable subtasks that help tackle longhorizon and/or sparse reward tasks. We propose a formalism for further abstracting the subtask structure by endowing an RM with the ability to call other RMs, thus composing a hierarchy of RMs (HRM) . We exploit HRMs by treating each call to an RM as an independently solvable subtask using the options framework, and describe a curriculum-based method to learn HRMs from traces observed by the agent. Our experiments reveal that exploiting a handcrafted HRM leads to faster convergence than with a flat HRM, and that learning an HRM is feasible in cases where its equivalent flat representation is not.  \n1. Introduction  \nMore than a decade ago, Dietterich et al. (2008) argued for the need to “learn at multiple time scales simultaneously, and with a rich structure of events and durations”. Finitestate machines (FSMs) are a simple yet powerful formalism for abstractly representing temporal tasks in a structured manner. One of the most prominent recent types of FSMs used in reinforcement learning (RL; Sutton & Barto, 2018) are reward machines (RMs; Toro Icarte et al., 2018 ; 2022), which compactly represent state-action histories in terms of high-level events; specifically, each edge is labeled with (i) a formula over a set of high-level events that capture a task’s subgoal, and (ii) a reward for satisfying the formula. Hence, RMs fulfill the need for structuring events and durations, and keep track of the achieved and pending subgoals.  \n1Imperial College London, UK 2ILASP Limited, UK 3Universitat Pompeu Fabra, Spain. Correspondence to: Daniel Furelos-Blanco \u003C[d.furelos-blanco18@imperial.ac.uk](d.furelos-blanco18@imperial.ac.uk)>.  \nProceedings of the 40 th International Conference on Machine Learning, Honolulu, Hawaii, USA. PMLR 202, 2023 . Copyright 2023 by the author(s) .  \nHierarchical reinforcement learning (HRL; Barto & Mahadevan, 2003) frameworks, such as options (Sutton et al., 1999), have been used to exploit RMs by learning policies at two levels of abstraction: (i) select a formula (i.e., subgoal) from a given RM state, and (ii) select an action to (eventually) satisfy the chosen formula (Toro Icarte et al., 2018 ; Furelos-Blanco et al., 2021) . The subtask decomposition powered by HRL enables learning at multiple scales simultaneously, and eases the handling of long-horizon and sparse reward tasks. In addition, several works have considered the problem of learning the RMs themselves from interaction (e.g., Toro Icarte et al., 2019 ; Xu et al., 2020 ; Furelos-Blanco et al., 2021 ; Hasanbeig et al., 2021) . A common problem among methods learning minimal RMs is that they scale poorly as the number of states grows.  \nIn this work, we make the following contributions:  \n1. Enhance the abstraction power of RMs by defining hierarchies of RMs (HRMs), where constituent RMs can call other RMs (Section 3) . We prove that any HRM can be transformed into an equivalent flat HRM that behaves exactly like the original RMs. We show that under certain conditions, the equivalent flat HRM can have exponentially more states and edges.  \n2. Propose an HRL algorithm to exploit HRMs by treating each call as a subtask (Section 4) . Learning policies in HRMs further fulfills the desiderata posed by Dietterich et al. since (i) there is an arbitrary number of time scales to learn across (not only two), and (ii) thereis a richer range of increasingly abstract events and durations. Besides, hierarchies enable modularity and, hence, the reusability of the RMs and policies. Empirically, we","cbCaiciH24y8A4S2","https://ap.wps.com/l/cbCaiciH24y8A4S2","pdf",2209459,1,48,"English","en",105,"# Introduction\n# Background\n## Reinforcement Learning Setup\n# Contributions\n## Hierarchies of RMs\n## HRL Algorithm for RM Calls\n## Curriculum-Based HRM Learning","[{\"question\":\"What problem do reward machines (RMs) address in reinforcement learning?\",\"answer\":\"RMs represent a task’s reward function using a finite-state machine whose edges encode subgoals as high-level events, enabling structured decomposition of long-horizon and sparse-reward tasks.\"},{\"question\":\"How do hierarchical reward machines (HRMs) extend reward machines?\",\"answer\":\"HRMs add the ability for one RM to call other RMs, creating a compositional hierarchy of subtasks rather than a single flat state machine.\"},{\"question\":\"How is HRM learning achieved from traces in this work?\",\"answer\":\"The method uses a curriculum-based approach that learns HRMs from traces for a set of composable tasks, leveraging simpler constituent RMs and previously learned components to make learning feasible.\"}]","Hierarchies of Reward Machines - curriculum-based hierarchical reward machines for reinforcement learning | PDF",1785810283,121,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"hierarchies-of-reward-machines-curriculum-based-hierarchical-reward-machines-for-reinforcement-learning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/technology/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/hierarchies-of-reward-machines-curriculum-based-hierarchical-reward-machines-for-reinforcement-learning/122371/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem do reward machines (RMs) address in reinforcement learning?","Question",{"text":75,"@type":76},"RMs represent a task’s reward function using a finite-state machine whose edges encode subgoals as high-level events, enabling structured decomposition of long-horizon and sparse-reward tasks.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do hierarchical reward machines (HRMs) extend reward machines?",{"text":80,"@type":76},"HRMs add the ability for one RM to call other RMs, creating a compositional hierarchy of subtasks rather than a single flat state machine.",{"name":82,"@type":73,"acceptedAnswer":83},"How is HRM learning achieved from traces in this work?",{"text":84,"@type":76},"The method uses a curriculum-based approach that learns HRMs from traces for a set of composable tasks, leveraging simpler constituent RMs and previously learned components to make learning feasible.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,113,118,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":111,"slug":112},50,"technology",{"id":114,"doc_module":4,"doc_module_name":46,"category_name":115,"show_sort_weight":116,"slug":117},7,"Healthcare",40,"healthcare",{"id":119,"doc_module":4,"doc_module_name":46,"category_name":120,"show_sort_weight":121,"slug":122},8,"Research & Report",30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]