[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84284-en":3,"doc-seo-84284-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84284,1374391974564,"Clementine","https://ap-avatar.wpscdn.com/avatar/14000253aa45c000a9e?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779874745381141002",8,"Research & Report","Reachability-Preserving Bellman Operator for the Discounted Reach-Cost Value Function","Hamilton–Jacobi (HJ) reachability offers rigorous safety and reachability guarantees for continuous-time dynamical systems, yet numerical computation suffers from the curse of dimensionality. Deep reinforcement learning (DRL) scales through sample-based learning, but standard RL uses additive cumulative rewards, while reachability is inherently non-additive. The work develops a semantics-preserving discounted reach-based value function and derives a non-additive Bellman operator whose unique fixed point matches the HJ formulation. Discounting renders the operator contractive, enabling existence, uniqueness, and convergence of value iteration, and provides a principled RL interpretation for solving the same fixed-point equation.","Reachability-Preserving Bellman Operator for the Discounted Reach-Cost Value Function: Uniting Hamilton–Jacobi Reachability and Reinforcement  \nLearning  \nIsabelle El-Hajj, Prashant Solanki, Jasper van Beers, Coen de Visser, and Erik-Jan van Kampen  \narXiv :2607 .07893v1 [ ee ss . SY] 8 Jul 2026  \nAbstract—Hamilton-Jacobi (HJ) reachability provides rigorous safety and reachability guarantees for continuous-time dynamical systems, but its numerical solution suffers from the curse of dimensionality. Deep reinforcement learning (DRL), by contrast, offers scalable sample-based methods. However, RL is typically built around additive cumulative rewards; whereas, reachability objectives are inherently non-additive. This mismatch makes a direct bridge between HJ reachability and RL nontrivial. Recent discounted formulations have either introduced contraction by altering the original reachability semantics, or preserved exact semantics on the HJ side without a corresponding Bellman fixedpoint characterization. In this paper, we close this gap by building on a semantics-preserving discounted reach-based value function and deriving a non-additive Bellman operator whose unique fixed point exactly matches the value function in the HJ formulation. We prove that discounting makes this operator contractive, yielding existence, uniqueness, and convergence of value iteration. Furthermore, we establish the equivalence between the HJ and Bellman characterizations, and show that RL can be interpreted as a sample-based approximation scheme for the same fixedpoint equation. This yields a principled and semantically exact connection between HJ reachability and RL, enabling learningbased methods to approximate reachability value functions while preserving their safety-critical meaning. As a result, the proposed framework opens the door to scalable, data-driven computation of reachable sets and safety certificates in high-dimensional systems. Numerical experiments demonstrate close agreement with Hamilton–Jacobi solutions, confirm preservation of reachability semantics via alignment of zero level sets, and support the interpretation of reinforcement learning as a sample-based solver of the proposed Bellman operator.  \nIndex Terms—Hamilton-Jacobi reachability, Non-additive Bellman Operator, Reinforcement learning, Reach cost, Safetycritical control.  \nI. INTRODUCTION  \nA. Background & Motivation  \nHamilton–Jacobi (HJ) reachability provides a principled framework for analyzing safety and reachability properties of continuous-time dynamical systems [18] . By characterizing reachable or safe sets as sublevel or superlevel sets of a value function satisfying a Hamilton–Jacobi variational inequality (HJVI), this framework offers strong semantic guarantees and has become a cornerstone of safety-critical control [3], [7] . However, its practical utility remains limited by the curse of dimensionality: numerical solutions of the associated partial  \nAll authors are with the section of Control & Simulation at the Faculty of Aerospace Engineering at Delft University of Technology.  \ndifferential equations scale exponentially with the state dimension [11], [12] .  \nA large body of work has therefore sought to improve the scalability of HJ reachability, including decomposition methods, level-set approximations, and learning-based surrogates [8], [14], [16], [17], [30] . Learning methods that are constrained by partial differential equation (PDEs), such as DeepReach [4], approximate the time-dependent HJ value function using a neural implicit representation trained to satisfy the PDE together with terminal and boundary conditions. While such approaches enforce local PDE consistency, they do not by themselves provide a global Bellman fixed-point characterization of the solution.  \nIn parallel, deep reinforcement learning (DRL) has emerged as a powerful paradigm for sequential decision-making and control in high-dimensional systems. Across games, robotics, fluid dynami","cbCaivccB8jaOwG7","https://ap.wps.com/l/cbCaivccB8jaOwG7","pdf",2896801,5,1,19,"English","en",105,"# Abstract\n# I. Introduction\n## A. Background & Motivation\n## B. Related Work & Contributions","[{\"question\":\"What core gap does the paper address between HJ reachability and reinforcement learning?\",\"answer\":\"It addresses the mismatch between HJ reachability objectives, which are inherently non-additive, and standard RL frameworks built on additive cumulative rewards, requiring an operator-theoretic bridge that preserves reachability semantics.\"},{\"question\":\"How does the paper construct a Bellman operator compatible with discounted reachability?\",\"answer\":\"It builds on a semantics-preserving discounted reach-based value function and derives a non-additive Bellman operator whose unique fixed point exactly matches the HJ value function.\"},{\"question\":\"What does discounting enable for the proposed operator?\",\"answer\":\"Discounting makes the operator contractive, leading to existence, uniqueness, and convergence of value iteration, and supporting a stable fixed-point solution perspective.\"}]",1784194587,48,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"reachability-preserving-bellman-operator-for-the-discounted-reach-cost-value-function","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/reachability-preserving-bellman-operator-for-the-discounted-reach-cost-value-function/84284/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What core gap does the paper address between HJ reachability and reinforcement learning?","Question",{"text":76,"@type":77},"It addresses the mismatch between HJ reachability objectives, which are inherently non-additive, and standard RL frameworks built on additive cumulative rewards, requiring an operator-theoretic bridge that preserves reachability semantics.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the paper construct a Bellman operator compatible with discounted reachability?",{"text":81,"@type":77},"It builds on a semantics-preserving discounted reach-based value function and derives a non-additive Bellman operator whose unique fixed point exactly matches the HJ value function.",{"name":83,"@type":74,"acceptedAnswer":84},"What does discounting enable for the proposed operator?",{"text":85,"@type":77},"Discounting makes the operator contractive, leading to existence, uniqueness, and convergence of value iteration, and supporting a stable fixed-point solution perspective.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":20,"slug":137},"General","general"]