[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84207-en":3,"doc-seo-84207-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84207,962075114765,"Quinn","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Safe Reinforcement Learning using Ideas from Model Predictive Control","Reinforcement learning (RL) can synthesize control policies from data, but real-world cyber-physical systems require strict, hard safety constraints during active learning and deployment. The proposal combines deep reinforcement learning’s adaptability with the formal safety guarantees of model predictive control (MPC). An offline MPC model computes a globally verified feasible state-action space that ensures constraint satisfaction. During training and deployment, the RL agent’s actions are deterministically projected by a safety filter onto this feasible set. Evaluation on a nonlinear 1-DOF hardware testbed shows safe exploration and stable policy convergence.","arXiv :2607 .07252v 1 [ cs .LG] 8 Jul 2026  \nSafe Reinforcement Learning using Ideas from Model Predictive Control  \nGeorg Sch¨afer 1 ,2 , Jakob Rehrl 1 , Stefan Huber 1 , and Simon Hirlaender2  \n1 Josef Ressel Centre for Intelligent and Secure Industrial Automation,  \nSalzburg University of Applied Sciences, Salzburg, Austria  \n2 Department of Artificial Intelligence and Human Interfaces,  \nParis Lodron University of Salzburg, Salzburg, Austria [georg.schaefer@fh-salzburg.ac.at](georg.schaefer@fh-salzburg.ac.at)  \nAbstract. Reinforcement learning (RL) enables the synthesis of control policies directly from data, making it highly appealing for complex cyber-physical systems (CPSs) and robotics. A persistent challenge, however, is ensuring strict, hard safety constraints during the active learning phase. In real-world physical systems, violating mechanical limits can cause irreversible damage, necessitating that exploration remains strictly within safe operational regions. We propose a generalized framework that combines the adaptive, high-performance nature of deep reinforcement learning (DRL) with the formal safety guarantees of model predictive control (MPC) . Using a mathematical model of the system dynamics, offline MPC computations define a feasible state-action space, representing all safe combinations of system states and control inputs that guarantee constraint satisfaction. During training and deployment, the RL agent’s instantaneous actions are projected onto this globally verified feasible set via a safety filter. We systematically evaluate our generalized approach on a non-linear 1-degree of freedom (1-DOF) laboratory testbed, demonstrating successful exploration and stable policy convergence on physical hardware.  \nKeywords: Safe reinforcement learning · Model predictive control · Cyber-physical systems · Control theory.  \n1 Introduction  \nThe application of reinforcement learning (RL) to cyber-physical systems (CPSs) and industrial robots holds immense promise. RL algorithms learn optimal control policies through active interaction with an environment, offering high potential to solve complex, non-linear control problems where traditional analytical models or linear approximations may fall short [16,10] . However, standard RL fundamentally relies on extensive, unconstrained exploration of the state-action space to discover these optimal policies.  \nThis core requirement creates a critical safety bottleneck: pure RL agents optimize for scalar reward signals but do not inherently respect strict physical  \n2 G. Sch¨afer et al.  \nsafety constraints [8] . In the context of real-world physical systems, such unconstrained exploration is unacceptable. Violating mechanical limits, thermal bounds, or actuator constraints can lead to hardware failure, costly downtime, or irreversible system damage. Therefore, an agent must avoid what we define as the unsafe state: any situation where the system either directly violates constraints or passes a “point of no return” where a future constraint violation becomes physically inevitable due to system dynamics, regardless of any subsequent control actions.  \nBridging the gap between the theoretical promise of RL and safe physical deployment requires mechanisms that guarantee absolute constraint satisfaction throughout the entire learning process. We propose a generalized framework combining the data-driven learning capabilities of RL with the formal, forwardlooking safety guarantees of offline model predictive control (MPC) . Rather than evaluating safety purely online, which is often computationally prohibitive for the high-frequency control loops required in modern CPSs, our approach uses offline MPC as an oracle to pre-define a “feasible state-action space”(F) . During training, a deterministic projection filter maps any potentially unsafe action proposed by the RL agent back into this safe set, avoiding violations during both exploration and deployment.  \n2 Related Work  \nThe cha","cbCaialvo5pBmvgF","https://ap.wps.com/l/cbCaialvo5pBmvgF","pdf",1615657,4,1,13,"English","en",105,"# Introduction\n# Related Work\n# Safe RL using Ideas from MPC","[{\"question\":\"Why is safety a bottleneck for applying reinforcement learning to physical systems?\",\"answer\":\"Standard RL relies on unconstrained exploration to discover optimal policies, which does not inherently respect strict physical safety constraints. In real hardware, violations of mechanical, thermal, or actuator limits can cause irreversible damage or failure.\"},{\"question\":\"How does the proposed framework use model predictive control to ensure safety?\",\"answer\":\"The method uses an offline MPC computation as an oracle to pre-define a feasible state-action space containing all safe state-input combinations. This space guarantees constraint satisfaction by construction.\"},{\"question\":\"How are unsafe actions handled during training and deployment?\",\"answer\":\"The RL agent may propose actions outside the safe set, and a deterministic safety filter projects such actions onto the globally verified feasible set, preventing violations during exploration and real-time deployment.\"}]",1784193954,33,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"safe-reinforcement-learning-using-ideas-from-model-predictive-control","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/safe-reinforcement-learning-using-ideas-from-model-predictive-control/84207/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is safety a bottleneck for applying reinforcement learning to physical systems?","Question",{"text":75,"@type":76},"Standard RL relies on unconstrained exploration to discover optimal policies, which does not inherently respect strict physical safety constraints. In real hardware, violations of mechanical, thermal, or actuator limits can cause irreversible damage or failure.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed framework use model predictive control to ensure safety?",{"text":80,"@type":76},"The method uses an offline MPC computation as an oracle to pre-define a feasible state-action space containing all safe state-input combinations. This space guarantees constraint satisfaction by construction.",{"name":82,"@type":73,"acceptedAnswer":83},"How are unsafe actions handled during training and deployment?",{"text":84,"@type":76},"The RL agent may propose actions outside the safe set, and a deterministic safety filter projects such actions onto the globally verified feasible set, preventing violations during exploration and real-time deployment.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]