[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86348-en":3,"doc-seo-86348-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86348,1099513958607,"Jiven","https://ap-avatar.wpscdn.com/avatar/100002390cf8733938c?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778829742770036399",8,"Research & Report","Disentangled Unsupervised Skill Discovery for Efficient Hierarchical Reinforcement Learning","A central challenge in unsupervised skill discovery is that learned skills often become entangled, where one latent skill variable simultaneously affects many state dimensions, making skill chaining and downstream reuse difficult. Disentangled Unsupervised Skill Discovery (DUSDi) learns disentangled skills by decomposing skill representations into components that each influence only a single state-space factor. Components can be concurrently composed for low-level action generation, then efficiently chained via hierarchical reinforcement learning to solve downstream tasks. DUSDi introduces a mutual-information-based objective to enforce disentanglement and uses value factorization for efficient optimization, outperforming prior methods across challenging environments.","Disentangled Unsupervised Skill Discovery for Efficient Hierarchical Reinforcement Learning  \nJiaheng Hu  \nUniversity of Texas at Austin [jiahengh@utexas.edu](jiahengh@utexas.edu)  \nPeter Stone†  \nUniversity of Texas at Austin, Sony AI[pstone@cs.utexas.edu](pstone@cs.utexas.edu)  \nZizhao Wang  \nUniversity of Texas at Austin [zizhao.wang@utexas.edu](zizhao.wang@utexas.edu)  \nRoberto Martín-Martín† University of Texas at Austin [robertomm@cs.utexas.edu](robertomm@cs.utexas.edu)  \narXiv :2410 . 1 125 1v2 [ cs .LG] 11 Jul 2026  \nAbstract  \nA hallmark of intelligent agents is the ability to learn reusable skills purely from unsupervised interaction with the environment. However, existing unsupervised skill discovery methods often learn entangled skills where one skill variable simultaneously influences many entities in the environment, making downstream skill chaining extremely challenging. We propose Disentangled Unsupervised Skill Discovery (DUSDi), a method for learning disentangled skills that can be efficiently reused to solve downstream tasks. DUSDi decomposes skills into disentangled components, where each skill component only affects one factor of the state space. Importantly, these skill components can be concurrently composed to generate low-level actions, and efficiently chained to tackle downstream tasks through hierarchical Reinforcement Learning. DUSDi defines a novel mutualinformation-based objective to enforce disentanglement between the influences of different skill components, and utilizes value factorization to optimize this objective efficiently. Evaluated in a set of challenging environments, DUSDi successfully learns disentangled skills, and significantly outperforms previous skill discovery methods when it comes to applying the learned skills to solve downstream tasks.  \nCode and skills visualization at [jiahenghu.github.io/DUSDi-site/](jiahenghu.github.io/DUSDi-site/) .  \n1 Introduction  \nReinforcement learning (RL) algorithms have achieved many successes in challenging tasks, including magnetic plasma control [11], automobile racing [54], and robotics [47] . However, applying existing RL algorithms to every new task in a tabula rasa manner often results in low sample efficiency that limits RL’s broader applicability [18] . Unsupervised skill discovery holds the promise of improving the sample efficiency of Reinforcement Learning, by learning a set of reusable skills through rewardfree interaction with the environment that can be later recombined to tackle multiple downstream tasks more efficiently. In practice, prior unsupervised RL skills are represented as a policy that conditions on a skill variable to generate diverse behaviors, and have led to successful and efficient learning of downstream tasks when combined with skill fine-tuning or hierarchical RL skill selection [13, 24, 58] .  \nDespite prior successes, a common limitation of the skills learned by existing unsupervised RL methods is that they are entangled: any change in the skill variable causes the agent to induce changes in multiple dimensions of the state space simultaneously. Learning to use and recombine these entangled skills can be extremely hard for an agent trying to solve downstream tasks, especially in  \n†Equal supervision.  \n38th Conference on Neural Information Processing Systems (NeurIPS 2024) .  \nPrior Works  \nDUSDi (ours)  \nFigure 1: Consider an agent practicing driving skills by learning to control a car’s speed (length of orange arrow), steering (curvature of orange arrow), and headlights (blue symbol),(Left) previous unsupervised skill discovery methods learn entangled skills, where a change in the skill variable can cause all three environment factors to change (Right) DUSDi learns disentangled skills with concurrent components, where each skill component only affects one factor of the state space, enabling efficient downstream task learning with hierarchical RL.  \ncomplex domains like multi-agent systems or household humanoid","cbCaidloatTOigp7","https://ap.wps.com/l/cbCaidloatTOigp7","pdf",2838997,4,1,18,"English","en",105,"# Introduction\n## Motivation and limitations of entangled skills\n## Key idea and state factorization\n## DUSDi approach overview\n## Intrinsic reward and objective design","[{\"question\":\"Why do existing unsupervised skill discovery methods struggle with downstream skill chaining?\",\"answer\":\"They often learn entangled skills, where changing one skill variable can affect multiple state dimensions at once, making it hard to reuse and recombine skills reliably.\"},{\"question\":\"What is DUSDi and what problem does it address?\",\"answer\":\"DUSDi (Disentangled Unsupervised Skill Discovery) learns disentangled skill components so that each component affects only one state factor, improving efficient reuse for downstream tasks.\"},{\"question\":\"How does DUSDi encourage disentanglement during learning?\",\"answer\":\"It defines a mutual-information-based intrinsic objective that rewards alignment between skill components and specific state factors while discouraging influence on other factors, optimized efficiently using value factorization.\"}]",1784210723,45,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"disentangled-unsupervised-skill-discovery-for-efficient-hierarchical-reinforcement-learning","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/disentangled-unsupervised-skill-discovery-for-efficient-hierarchical-reinforcement-learning/86348/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do existing unsupervised skill discovery methods struggle with downstream skill chaining?","Question",{"text":75,"@type":76},"They often learn entangled skills, where changing one skill variable can affect multiple state dimensions at once, making it hard to reuse and recombine skills reliably.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is DUSDi and what problem does it address?",{"text":80,"@type":76},"DUSDi (Disentangled Unsupervised Skill Discovery) learns disentangled skill components so that each component affects only one state factor, improving efficient reuse for downstream tasks.",{"name":82,"@type":73,"acceptedAnswer":83},"How does DUSDi encourage disentanglement during learning?",{"text":84,"@type":76},"It defines a mutual-information-based intrinsic objective that rewards alignment between skill components and specific state factors while discouraging influence on other factors, optimized efficiently using value factorization.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]