[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-116859-en":3,"doc-seo-116859-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},116859,1099513958762,"Logic","https://ap-avatar.wpscdn.com/avatar/1000023916a998db790?x-image-process=image/resize,m_fixed,w_180,h_180&k=1784791008015729253",8,"Research & Report","Deep Black-Box Reinforcement Learning with Movement Primitives","Episode-based reinforcement learning treats reinforcement learning as a black-box optimization problem, learning controller parameters such as movement primitives conditioned on a context descriptor. The approach enables smooth control trajectories, supports non-Markovian reward definitions, and performs exploration directly in movement-primitive parameter space, making it attractive for sparse-reward tasks. High-dimensional primitive parameters have limited deep learning usage, motivating a new deep episode-based RL algorithm.","Deep Black-Box Reinforcement Learning with Movement Primitives  \nFabian Otto 1 ;2 , Onur Celik3 , Hongyi Zhou3 , Hanna Ziesche 1 , Ngo Anh Vien 1 ,  \nand Gerhard Neumann3  \n1 Bosch Center for Artiﬁcial Intelligence, Germany  \n2 University of T¨ubingen, Germany  \n3 Autonomous Learning Robots, Karlsruhe Institute of Technology, Germany  \n[fabian.otto@bosch.com](fabian.otto@bosch.com)  \nAbstract: Episode-based reinforcement learning (ERL) algorithms treat reinforcement learning (RL) as a black-box optimization problem where we learn to select a parameter vector of a controller, often represented as a movement primitive, for a given task descriptor called a context. ERL offers several distinct beneﬁts in comparison to step-based RL. It generates smooth control trajectories, can handle non-Markovian reward deﬁnitions, and the resulting exploration in parameter space is well suited for solving sparse reward settings. Yet, the high dimensionality of the movement primitive parameters has so far hampered the effective use of deep RL methods. In this paper, we present a new algorithm for deep ERL.  \nIt is based on differentiable trust region layers, a successful on-policy deep RL algorithm. These layers allow us to specify trust regions for the policy update that are solved exactly for each state using convex optimization, which enables policies learning with the high precision required for the ERL. We compare our ERL algorithm to state-of-the-art step-based algorithms in many complex simulated robotic control tasks. In doing so, we investigate different reward formulations  \n-dense, sparse, and non-Markovian. While step-based algorithms perform well only on dense rewards, ERL performs favorably on sparse and non-Markovian rewards. Moreover, our results show that the sparse and the non-Markovian rewards are also often better suited to deﬁne the desired behavior, allowing us to obtain considerably higher quality policies compared to step-based RL.  \nKeywords: Movement Primitives, Deep Episode-Based RL, Trust Region Layers  \n1 Introduction  \nReinforcement learning (RL) problems can be viewed from a step-based [1, 2, 3] and an episodebased perspective [4, 5, 6] . In the former, most commonly found in deep RL, a policy selectsan action for each state of the trajectory. Consequently, step-based RL requires Markovian reward deﬁnitions. Moreover, the exploration in action space often induces a random walk that inadequately explores the entire behavior space of the agent. As a result, most step-based RL algorithms only work well with dense rewards, where the agent receives a reward signal at each time step, rather than only at the last step. We refer to the latter as a sparse reward setting.  \nIn the episode-based reinforcement learning (ERL) perspective, we choose the complete behavior during the episode with respect to the controller parameters in the beginning of the episode. Typically, simple controllers are used that are valid only for executing a single trajectory, such as movement primitives (MPs) . At the beginning of an episode, the MP parameters are adapted to a context vector that serves as task descriptor and may contain e. g. different initial joint conﬁgurations, the goal to reach, or start positions as well as object and obstacles locations. ERL allows for efﬁcient exploration of the behavior space because exploration is implemented in the parameter space of the MP, allowing these algorithms to learn from sparse or even non-Markovian rewards. In addition, ERL inherently generates smooth control trajectories due to the use of MPs, which often leads to more energy efﬁcient behavior. However, ERL has so far not gained much popularity as  \n6th Conference on Robot Learning (CoRL 2022), Auckland, New Zealand.  \nfor t in 1; 2; : : : ; T  \nContext  \nc  \nPolicy 􀀙􀀒 (w jc)  \nDesired Trajectory  \n􀀜 d = 􀀉(w) = (sd1 ; sd2 ; : : : ; sdT)  \nController f (st ; sdt)  \nat  \nst+1  \nEnvironment  \nFigure 1: Overview of the proposed Black-Box Reinforce","cbCaidjip6LPDCM2","https://ap.wps.com/l/cbCaidjip6LPDCM2","pdf",2565110,1,22,"English","en",105,"# Introduction\n## Episode-based vs step-based reinforcement learning\n## Movement primitives and context adaptation\n# Proposed deep ERL method\n## Differentiable trust region layers and exact policy updates\n## Reward formulations and comparative evaluation\n# Experiments and results\n## Dense, sparse, and non-Markovian reward settings","[{\"question\":\"What is episode-based reinforcement learning (ERL) in this work?\",\"answer\":\"ERL frames reinforcement learning as black-box optimization over controller parameters at the start of each episode, conditioned on a context vector rather than selecting actions at every step.\"},{\"question\":\"Why are movement primitives useful for ERL?\",\"answer\":\"Movement primitives parameterize a whole trajectory and yield smooth control behavior; adapting primitive parameters to context enables efficient exploration in parameter space and supports sparse or non-Markovian rewards.\"},{\"question\":\"How does the proposed deep ERL method differ from prior deep step-based RL?\",\"answer\":\"It adapts differentiable trust region projection layers to the ERL setting, providing exact trust regions via convex optimization per state, which improves precision for learning in high-dimensional parameter spaces.\"}]","Deep Black-Box Reinforcement Learning with Movement Primitives | PDF",1785672110,55,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"deep-black-box-reinforcement-learning-with-movement-primitives","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/deep-black-box-reinforcement-learning-with-movement-primitives/116859/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is episode-based reinforcement learning (ERL) in this work?","Question",{"text":75,"@type":76},"ERL frames reinforcement learning as black-box optimization over controller parameters at the start of each episode, conditioned on a context vector rather than selecting actions at every step.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why are movement primitives useful for ERL?",{"text":80,"@type":76},"Movement primitives parameterize a whole trajectory and yield smooth control behavior; adapting primitive parameters to context enables efficient exploration in parameter space and supports sparse or non-Markovian rewards.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the proposed deep ERL method differ from prior deep step-based RL?",{"text":84,"@type":76},"It adapts differentiable trust region projection layers to the ERL setting, providing exact trust regions via convex optimization per state, which improves precision for learning in high-dimensional parameter spaces.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]