[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-128112-en":3,"doc-seo-128112-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},128112,3985741905716,"Rowan","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","No-Regret Reinforcement Learning in Smooth MDPs - Paper Abstract","Obtaining no-regret guarantees for reinforcement learning with continuous state and/or action spaces remains a central open problem. Existing approaches work only in narrowly tailored settings, leaving the general case unresolved. This paper proposes a new structural assumption for smooth Markov decision processes and introduces two algorithms based on an orthogonal feature map using Legendre polynomials. LEGENDREELEANOR offers no-regret under weaker conditions but is computationally inefficient, while LEGENDRE-LSVI achieves polynomial-time performance for a smaller problem class and delivers the strongest regret bounds.","No-Regret Reinforcement Learning in Smooth MDPs  \nDavide Maran 1 Alberto Maria Metelli 1 Matteo Papini 1 Marcello Restelli 1  \nAbstract  \nObtaining no-regret guarantees for reinforcement learning (RL) in the case of problems with continuous state and/or action spaces is still one of the major open challenges in the field. Recently, a variety of solutions have been proposed, but besides very specific settings, the general problem remains unsolved. In this paper, we introduce a novel structural assumption on the Markov decisiothapgceerasesizes(MDPsmost o)f, nthaeeeytiνngsmproopothnosedessso, far (e.g., linear MDPs and Lipschitz MDPs) . To face this challenging scenario, we propose two algoMriDthPsms fBoogretalgorminiithmizationbuilduinpoνnshoite of constructing an MDP representation through an orthogonal feature map based on Legendre polynomials. The first algorithm, LEGENDREELEANOR, archives the no-regret property under weaker assumptions but is computationally inefficient, whereas the second one, LEGENDRE-LSVI, runs in polynomial time, although for a smaller class of problems. After analyzing their regret properties, we compare our results with state-ofthe-art ones from RL theory, showing that our algorithms achieve the best guarantees.  \n1. Introduction  \nReinforcement learning (RL) (Sutton & Barto, 2018) is a paradigm of artificial intelligence in which the agent interacts with an environment to maximize a reward signal in the long term. From the theoretical perspective, a lot of effort has been put into designing algorithms with small (cumulative) regret, which is an index of how much the policies (i.e., the behavior) played by the algorithm during the learning process are suboptimal. For the case of tabular Markov decision processes (MDPs), an optimal result was  \n*Equal contribution 1Politecnico di Milano, Milan, Italy. Correspondence to: Davide Maran \u003C[davide.maran@polimi.it](davide.maran@polimi.it) >.  \nProceedings of the 41 st International Conference on Machine Learning, Vienna, Austria. PMLR 235, 2024 . Copyright 2024 by the author(s) .  \nfirst proved by Azar et al. (2017), who showed a bound osoe,,t ofAanisdorda fiHetrnhpii| ip|z| |Keo, , heKryreisepSthisfinmbThi regret is minimax-optimal, in the sense that no algorithm can achieve smaller regret for every arbitrary tabular MDP. Unfortunately, assuming that the state-action space is finite is extremely restrictive, as the number of states and/or actions can be huge or even infinite in practice. This is especially critical for a large variety of real-world scenarios in which RL has achieved successful results, including robotics (Kober et al., 2013), autonomous driving (Kiranet al., 2021), and trading (Hambly et al., 2023) . These scenarios are usually modeled as MDPs with continuous state and/or action spaces, as the underlying dynamics is too complex to be captured by a finite number of states and/or actions. It is not by chance that one of the most common benchmarks for RL algorithms, MUJOCO (Todorov et al., 2012 ; Brockman et al., 2016), is composed of environments characterized by continuous state and action spaces. This highlights the notable gap between the current maturity of theory and the pressing needs of practical application. For this reason, devising algorithms with regret bounds for RL in continuous spaces is currently one of the most important challenges of the whole field.  \nSince, without any further assumption, the RL problem in continuous spaces is non-learnable, 1 the modern literature revolves around searching for the weakest structural assumptions under which the problem can be solved efficiently. Linear quadratic regulator (LQR) (Bemporad et al., 2002) is a model for the environment that is widely used in control theory, where the state of the system evolves according to a linear dynamical system and the reward is quadratic. For the online control of this problem, when  the sywaseceme obmpxdnibalunkAbinnbefn-Yen,grkolgibou& Sthmnzd oepeThrd´ar","cbCaioBP34cvwMjb","https://ap.wps.com/l/cbCaioBP34cvwMjb","pdf",527428,2,1,30,"English","en",105,"# Introduction\n## Continuous-state reinforcement learning challenge\n## Structural assumptions: linear and Lipschitz MDPs\n## Proposed smooth-MDP framework and algorithms","[{\"question\":\"What challenge does the paper target in reinforcement learning theory?\",\"answer\":\"It targets obtaining no-regret guarantees for reinforcement learning when state and/or action spaces are continuous, a problem that is unsolved in general settings.\"},{\"question\":\"What structural assumption does the paper introduce?\",\"answer\":\"The paper introduces a novel structural assumption on smooth Markov decision processes to make regret analysis possible beyond specific linear or Lipschitz cases.\"},{\"question\":\"How do LEGENDREELEANOR and LEGENDRE-LSVI differ?\",\"answer\":\"LEGENDREELEANOR provides no-regret under weaker assumptions but is computationally inefficient, while LEGENDRE-LSVI runs in polynomial time but applies to a smaller class of problems.\"}]","No-Regret Reinforcement Learning in Smooth MDPs - Paper Abstract | PDF",1785944891,76,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"no-regret-reinforcement-learning-in-smooth-mdps-paper-abstract","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,48,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":20},"https://docshare.wps.com/document/","Document",{"item":49,"name":12,"@type":44,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/no-regret-reinforcement-learning-in-smooth-mdps-paper-abstract/128112/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-23","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What challenge does the paper target in reinforcement learning theory?","Question",{"text":76,"@type":77},"It targets obtaining no-regret guarantees for reinforcement learning when state and/or action spaces are continuous, a problem that is unsolved in general settings.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What structural assumption does the paper introduce?",{"text":81,"@type":77},"The paper introduces a novel structural assumption on smooth Markov decision processes to make regret analysis possible beyond specific linear or Lipschitz cases.",{"name":83,"@type":74,"acceptedAnswer":84},"How do LEGENDREELEANOR and LEGENDRE-LSVI differ?",{"text":85,"@type":77},"LEGENDREELEANOR provides no-regret under weaker assumptions but is computationally inefficient, while LEGENDRE-LSVI runs in polynomial time but applies to a smaller class of problems.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":22,"slug":122},"research-report",{"id":124,"doc_module":4,"doc_module_name":47,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":47,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":47,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":47,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]