[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85392-en":3,"doc-seo-85392-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85392,1099513958607,"Jiven","https://ap-avatar.wpscdn.com/avatar/100002390cf8733938c?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778829742770036399",8,"Research & Report","Distributionally Robust Reinforcement Learning with Interactive Data Collection","The sim-to-real gap between training and testing environments makes reinforcement learning hard. Distributionally robust RL, cast as a robust Markov decision process, seeks a policy that performs well in the worst case over a pre-specified uncertainty set around the training environment. This work studies robust RL under interactive data collection, where the learner refines the policy from trial-and-error. It proves sample-efficiency is unattainable due to the curse of support shift, then introduces a vanishing minimal value assumption for TV-robust sets to eliminate support-shift pathologies and enables near-optimal algorithms. The framework is extended to robust set variants and Markov games, and applied to data-driven robust inventory control with explicit learning guarantees under demand shifts.","arXiv :2404 .03578v 3 [ cs .LG] 13 Jul 2026  \nDistributionally Robust Reinforcement Learning with Interactive Data Collection: Fundamental Hardness and Near-Optimal  \nAlgorithms  \nMiao Lu∗† Han Zhong∗‡ Tong Zhang§ Jose Blanchet†  \nApril 5, 2024; Revised: July 13, 2026  \nAbstract  \nThe sim-to-real gap, which represents the disparity between training and testing environments, poses a significant challenge in reinforcement learning (RL) . A promising approach to addressing this challenge is distributionally robust RL, often framed as a robust Markov decision process (RMDP) . In this framework, the objective is to find a robust policy that achieves good performance under the worst-case scenario among all environments within a pre-specified uncertainty set centered around the training environment. Unlike previous work, which relies on a generative model or a pre-collected offline dataset enjoying good coverage of the deployment environment, we tackle robust RL via interactive data collection, where the learner interacts with the training environment only and refines the policy through trial and error. In this robust RL paradigm, two main challenges emerge: managing distributional robustness while striking a balance between exploration and exploitation during data collection. Initially, we establish that sampleefficient learning without additional assumptions is unattainable owing to the curse of support shift; i.e. , the potential disjointedness of the distributional supports between the training and testing environments. To circumvent such a hardness result, we introduce the vanishing minimal value assumption to RMDPs with a total-variation (TV) distance robust set, postulating that the minimal value of the optimal robust value function is zero. We prove that such an assumption effectively eliminates support shift pathologies for RMDPs with a TV distance robust set, and present an algorithm with near-optimal sample complexity. To demonstrate the breadth of our framework, we further extend our algorithm and theory to new robust set formulations and robust Markov game settings. Finally, to illustrate the operational relevance of our framework, we apply our algorithm to the data-driven robust inventory control, yielding explicit learning guarantees for robust decision-making under demand shifts. Our work makes the initial step to uncovering the inherent difficulty of robust RL via interactive data collection and sufficient conditions for designing a sample-efficient algorithm accompanied by sharp sample complexity analysis.  \nKeywords: distributionally robust reinforcement learning, interactive data collection, robust Markov decision process, robust Markov game, sample complexity, online regret  \n∗ Equal contributions. Email [to](to miaolu@stanford.edu)[ miaolu@stanford.edu](to miaolu@stanford.edu), [hanzhong@stu.pku.edu.cn](hanzhong@stu.pku.edu.cn)[ ](hanzhong@stu.pku.edu.cn)†Department of Management Science and Engineering, Stanford University.‡Center for Data Science, Peking University.  \n§ Department of Computer Science, University of Illinois Urbana-Champaign.  \nContents  \n1 Introduction 4  \n1.1 Contributions ............................................. 5  \n1.2 Related Works ............................................ 7  \n1.3 Notations ............................................... 9  \n2 Preliminaries 9  \n2.1 Robust Markov Decision Processes ................................. 9  \n2.2 Robust RL with Interactive Data Collection ............................ 12  \n3 A Hardness Result: The Curse of Support Shift 12  \n4 A Solvable Case, Efficient Algorithm, and Sharp Analysis 14  \n4.1 Vanishing Minimal Value: Eliminating Support Shift ....................... 14  \n4.2 Algorithm Design: OPROVI-TV .................................. 16  \n4.2.1 Training Environment Transition Estimation ....................... 16  \n4.2.2 Optimistic Robust Planning ................................. 17  \n4.3 Theoretical Guarantees ..........................","cbCaipYAryDhlpto","https://ap.wps.com/l/cbCaipYAryDhlpto","pdf",815068,6,1,61,"English","en",105,"# Introduction\n## Contributions\n# Preliminaries\n## Robust Markov Decision Processes\n# A Hardness Result: The Curse of Support Shift\n# A Solvable Case, Efficient Algorithm, and Sharp Analysis\n## Vanishing Minimal Value: Eliminating Support Shift\n# Extension I: Robust Set with Bounded Transition Probability Ratio\n# Extension II: Robust Decision Making in Multi-Agent Systems\n## Algorithm and Theory\n# Application: Data-Driven Robust Inventory Control\n## Inventory Control as a Finite-horizon MDP\n# Conclusions and Discussions","[{\"question\":\"What problem does this paper target in reinforcement learning?\",\"answer\":\"It targets the sim-to-real gap, where training and testing environments differ, making learned policies unreliable without robust treatment.\"},{\"question\":\"How is distributional robustness formulated in this work?\",\"answer\":\"It is framed as a robust Markov decision process, optimizing policy performance under the worst case within an uncertainty set around the training environment.\"},{\"question\":\"Why is sample-efficient learning impossible in general for interactive robust RL?\",\"answer\":\"The paper shows an inherent hardness from the curse of support shift, meaning training and testing distributions may have disjoint support, preventing efficient learning without extra assumptions.\"}]",1784203098,154,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"distributionally-robust-reinforcement-learning-with-interactive-data-collection","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/distributionally-robust-reinforcement-learning-with-interactive-data-collection/85392/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does this paper target in reinforcement learning?","Question",{"text":76,"@type":77},"It targets the sim-to-real gap, where training and testing environments differ, making learned policies unreliable without robust treatment.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How is distributional robustness formulated in this work?",{"text":81,"@type":77},"It is framed as a robust Markov decision process, optimizing policy performance under the worst case within an uncertainty set around the training environment.",{"name":83,"@type":74,"acceptedAnswer":84},"Why is sample-efficient learning impossible in general for interactive robust RL?",{"text":85,"@type":77},"The paper shows an inherent hardness from the curse of support shift, meaning training and testing distributions may have disjoint support, preventing efficient learning without extra assumptions.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]