[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121983-en":3,"doc-seo-121983-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},121983,4398048950312,"Violet","https://ap-avatar.wpscdn.com/avatar/400002538284de19e3c?_k=1778320343897328908",8,"Research & Report","Improving Single and Multi-Agent Deep Reinforcement Learning Methods - Thesis abstract","Reinforcement Learning (RL) trains an agent to make decisions via rewards or penalties from environment interactions, and Deep RL combines RL with deep neural networks to handle complex, high-dimensional data for long-horizon sequential decisions. This thesis examines obstacles that hinder learning in specific environments and proposes methods to improve performance, sample efficiency, and generalizability. For single-agent RL, it addresses sparse-reward exploration using semantic exploration. For cooperative multi-agent RL, it tackles coordination and joint-action exploration with universal value exploration, scalable role-based learning, and analysis of independent policy methods.","Improving Single and Multi-Agent Deep Reinforcement Learning Methods  \nTarun Gupta  \nExeter College  \nUniversity of Oxford  \nA thesis submitted for the degree of Doctor of Philosophy  \nMichaelmas Term 2023  \n2  \nAbstract  \nReinforcement Learning (RL) is a framework where an agent learns to make decisions using data-driven feedback from interactions with the environment in the form of rewards or penalties for actions. Deep RL integrates deep learning with RL, harnessing the power of deep neural networks to process complex, high-dimensional data. Using the framework of deep RL, our machine learning research community has achieved tremendous progress in enabling machines to make sequential decisions over long time horizons. These advances include attaining super-human performance in Atari [Mnih et al., 2015], mastering the game of Go, beating the human worldchampion [Silver et al., 2017], providing robust recommendation systems [GomezUribe and Hunt, 2015, Singh et al., 2021] . This thesis focuses on identifying some key challenges that impede the learning of RL agents within their specific environmentsand improving the methods leading to better performance of agents, improved sample efficiency, and generalizability of learned agent policies.  \nIn Part I of the thesis, we focus on exploration in single-agent RL settings where an agent must interact with a complex environment to pursue a goal. An agent that fails to explore its environment is unlikely to achieve high performance, as it will miss critical rewards and, as a result, cannot learn optimal behavior. One key challenge arises from sparse reward environments where the agent only receives feedback once the task is completed, making exploration more challenging. We propose a novel method that enables semantic exploration, resulting in higher sample efficiency and performance on sparse reward tasks.  \nIn Part II of the thesis, we focus on cooperative Multi-Agent Reinforcement Learning (MARL), an extension of the usual RL setting, where we consider multiple agents interacting in the same environment toward a shared task. In multi-agent tasks requiring significant coordination among agents with strict penalties formiscoordination, state-of-the-art MARL methods often fail to learn useful behaviors as agents get stuck in a sub-optimal equilibrium. Another challenge is exploration in the joint action space of all agents, which grows exponentially with the number of agents. To address these challenges, we propose innovative approaches like universal value exploration and scalable role-based learning. These methods facilitate improved coordination among agents, faster exploration, and enhance the agents’ ability to adapt to new environments and tasks, showcasing zero-shot generalization capabilities and resulting in higher sample efficiency. Lastly, we investigate independent policybased methods in cooperative MARL, where each agent considers other agents as part of the environment. We show that such methods can perform better than state-of-the-art joint learning approaches on a popular multi-agent benchmark.  \nIn summary, the contributions of this thesis significantly improve the stateof-the-art in deep (multi-agent) reinforcement learning. The agents developed in this thesis can explore their environments efficiently to improve sample efficiency, learn tasks that require significant multi-agent coordination, and enable zero-shot generalization across various tasks.  \nImproving Single and Multi-Agent Deep Reinforcement Learning Methods  \nTarun Gupta  \nExeter College University of Oxford  \nA thesis submitted for the degree of Doctor of Philosophy  \nMichaelmas Term 2023  \nDedicated to my family  \nAcknowledgements  \nI am grateful to my supervisor, Shimon Whiteson for his unwavering support, guidance, and encouragement throughout my PhD journey. His insights and expertise have been invaluable, and I feel fortunate to have had the opportunity to work under his mentorship. I extend my hea","cbCaikjJD4s7Itb8","https://ap.wps.com/l/cbCaikjJD4s7Itb8","pdf",17703140,1,219,"English","en",105,"# Abstract\n## Single-agent exploration\n## Cooperative multi-agent reinforcement learning\n## Summary of contributions","[{\"question\":\"What is the core focus of this thesis in deep reinforcement learning?\",\"answer\":\"The thesis identifies challenges that limit RL agent learning in their environments and develops methods to improve agent performance, sample efficiency, and generalizability.\"},{\"question\":\"How does the thesis address exploration in single-agent RL?\",\"answer\":\"It targets sparse-reward settings where feedback appears only after task completion, proposing a semantic exploration method to improve sample efficiency and performance.\"},{\"question\":\"What challenges arise in cooperative multi-agent RL, and how are they addressed?\",\"answer\":\"The thesis highlights coordination difficulties and exploration in the exponentially growing joint action space. It proposes approaches such as universal value exploration and scalable role-based learning to enhance coordination, exploration, and adaptation, including zero-shot generalization.\"}]","Improving Single and Multi-Agent Deep Reinforcement Learning Methods - Thesis abstract | PDF",1785808146,552,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"improving-single-and-multi-agent-deep-reinforcement-learning-methods-thesis-abstract","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/improving-single-and-multi-agent-deep-reinforcement-learning-methods-thesis-abstract/121983/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the core focus of this thesis in deep reinforcement learning?","Question",{"text":75,"@type":76},"The thesis identifies challenges that limit RL agent learning in their environments and develops methods to improve agent performance, sample efficiency, and generalizability.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the thesis address exploration in single-agent RL?",{"text":80,"@type":76},"It targets sparse-reward settings where feedback appears only after task completion, proposing a semantic exploration method to improve sample efficiency and performance.",{"name":82,"@type":73,"acceptedAnswer":83},"What challenges arise in cooperative multi-agent RL, and how are they addressed?",{"text":84,"@type":76},"The thesis highlights coordination difficulties and exploration in the exponentially growing joint action space. It proposes approaches such as universal value exploration and scalable role-based learning to enhance coordination, exploration, and adaptation, including zero-shot generalization.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]