[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117474-en":3,"doc-seo-117474-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117474,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","Causal Reinforcement Learning - A Survey","Reinforcement learning offers a core framework for sequential decision-making under uncertainty, yet real-world deployment remains difficult due to limited world understanding, reliance on extensive trial-and-error interaction, and gaps in explanation and knowledge generalization. Causality strengthens this paradigm by systematizing knowledge and enabling invariance-driven transfer. This survey reviews causal reinforcement learning, detailing how causal relationships enhance sample efficiency, generalizability, knowledge transfer, spurious-correlation mitigation, explainability, fairness, and safety, then discusses current limitations and future research directions.","Causal Reinforcement Learning: A Survey  \nZhihong Deng [zhi-hong.deng@student.uts. edu. au](zhi-hong.deng@student.uts. edu. au)  \nAustralian Artificial Intelligence Institute, University of Technology Sydney  \nJing Jiang [jing.jiang@uts. edu. au](jing.jiang@uts. edu. au)  \nAustralian Artificial Intelligence Institute, University of Technology Sydney  \nGuodong Long [guodong.long@uts. edu. auu](guodong.long@uts. edu. auu)  \nAustralian Artificial Intelligence Institute, University of Technology Sydney  \nChengqi Zhang [chengqi.zhang@uts. edu. au](chengqi.zhang@uts. edu. au)  \nAustralian Artificial Intelligence Institute, University of Technology Sydney  \nReviewed on OpenReview: [https: // openreview. net/ forum? id= qqnt tX9LPo](https: // openreview. net/ forum? id= qqnt tX9LPo)  \nAbstract  \nReinforcement learning is an essential paradigm for solving sequential decision problems under uncertainty. Despite many remarkable achievements in recent decades, applying reinforcement learning methods in the real world remains challenging. One of the main obstacles is that reinforcement learning agents lack a fundamental understanding of the world and must therefore learn from scratch through numerous trial-and-error interactions. They may also face challenges in providing explanations for their decisions and generalizing the acquired knowledge. Causality, however, offers notable advantages by formalizing knowledge in a systematic manner and harnessing invariance for effective knowledge transfer. This has led to the emergence of causal reinforcement learning, a subfield of reinforcement learning that seeks to enhance existing algorithms by incorporating causal relationships into the learning process. In this survey, we provide a comprehensive review of the literature in this domain. We begin by introducing basic concepts in causality and reinforcement learning, and then explain how causality can help address key challenges faced by traditional reinforcement learning. We categorize and systematically evaluate existing causal reinforcement learning approaches, with a focus on their ability to enhance sample efficiency, advance generalizability, facilitate knowledge transfer, mitigate spurious correlations, and promote explainability, fairness, and safety. Lastly, we outline the limitations of current research and shed light on future directions in this rapidly evolving field.  \n1 Introduction  \n“All reasonings concerning matter of fact seem to be founded on the relation of cause and effect. By means of that relation alone we can go beyond the evidence of our memory and senses . \"  \n—David Hume, An Enquiry Concerning Human Understanding.  \nHumans possess an inherent capacity to grasp the concept of causality from a young age (Wellman, 1992; Inagaki & Hatano, 1993; Koslowski & Masnick, 2002; Sobel & Sommerville, 2010) . This innate understanding empowers us to recognize that altering specific factors can lead to corresponding outcomes, enabling us to actively manipulate our surroundings to accomplish desired objectives and acquire fresh insights. A deep understanding of cause and effect enables us to explain behaviors (Schult & Wellman, 1997), predict future outcomes (Shultz, 1982), and use counterfactual reasoning to dissect past events (Harris et al., 1996) . These cognitive abilities inherently shape human thought and reasoning (Sloman, 2005; Sloman & Lagnado, 2015;  \n\n|  |  |\n| --- | --- |\n|  | Does changing the size of an object affect the outcome? What about shape and color? Are there any invariants across tasks and environments? |\n\n\n|  | \u003Cbr>What is the effect of certain treatments? Does receiving the treatment directly contribute to the patient's recovery, or are other unmeasured factors at play? |\n| --- | --- |\n\nFigure 1: Illustrative examples of causality in reinforcement learning and its impact on decision making.  \nPearl, 2009b; Pearl & Mackenzie, 2018), forming the basis for modern society and civilization, as well as propelling ad","cbCaiox4pF79DfkI","https://ap.wps.com/l/cbCaiox4pF79DfkI","pdf",1853948,1,52,"English","en",105,"# Introduction\n## Causality in reinforcement learning applications\n### Robotic manipulation and invariance\n### Medical scenarios and confounders\n### Autonomous driving and counterfactual reasoning","[{\"question\":\"为什么传统强化学习在现实中更难落地？\",\"answer\":\"强化学习智能体缺乏对世界的基本理解，通常需要通过大量试错交互从零学习；此外还会面临解释决策、以及将学到的知识推广到新情境的困难。\"},{\"question\":\"因果方法如何帮助提升强化学习？\",\"answer\":\"因果建模把知识以系统方式形式化，并利用不变性促进有效的知识迁移，从而增强学习过程中的关键能力。\"},{\"question\":\"调查重点覆盖哪些改进方向？\",\"answer\":\"文中系统梳理现有因果强化学习方法，重点评估其在样本效率、可泛化性、知识迁移、减轻虚假相关、解释性、公平性与安全性方面的作用，同时总结研究不足并展望未来方向。\"}]","Causal Reinforcement Learning - A Survey | PDF",1785676059,131,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"causal-reinforcement-learning-a-survey","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/causal-reinforcement-learning-a-survey/117474/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"为什么传统强化学习在现实中更难落地？","Question",{"text":75,"@type":76},"强化学习智能体缺乏对世界的基本理解，通常需要通过大量试错交互从零学习；此外还会面临解释决策、以及将学到的知识推广到新情境的困难。","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"因果方法如何帮助提升强化学习？",{"text":80,"@type":76},"因果建模把知识以系统方式形式化，并利用不变性促进有效的知识迁移，从而增强学习过程中的关键能力。",{"name":82,"@type":73,"acceptedAnswer":83},"调查重点覆盖哪些改进方向？",{"text":84,"@type":76},"文中系统梳理现有因果强化学习方法，重点评估其在样本效率、可泛化性、知识迁移、减轻虚假相关、解释性、公平性与安全性方面的作用，同时总结研究不足并展望未来方向。","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]