[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85060-en":3,"doc-seo-85060-105":29,"detail-sidebar-cat-0-en-105":82},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85060,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","Data-Driven Critic-Free Policy Iteration for Continuous-Time Linear Quadratic Regulation","Data-driven off-policy policy iteration for continuous-time linear quadratic regulation is reformulated to remove explicit critic identification. The method anchors the Riccati equation at a known stabilizing gain and rewrites optimality as a policy-space residual, then uses an endpoint null-space projection to eliminate the value-matrix term from the integral data equation. Under a projected rank condition, the data equation matches the policy-space residual and each least-squares update coincides with the Kleinman iteration, preserving stability and convergence while reducing regression dimensions and rank assumptions.","Data-Driven Critic-Free Policy Iteration for Continuous-Time Linear Quadratic Regulation  \nJiacheng Wu, Yang Zhu∗ , Hongye Su, Senior Member, IEEE  \narXiv :2607 .08204v1 [ ee ss . SY] 9 Jul 2026  \nAbstract—For continuous-time linear quadratic regulation with unknown system matrices, data-driven off-policy policy iteration typically estimates the value matrix and the improved feedback gain through a joint critic–actor regression. We show that the critic is not needed in the policy-improvement step. The key is to anchor the Riccati equation at a known stabilizing gain and express optimality as a policy-space residual. An endpoint null-space projection then removes the value-matrix term from the integral data equation. This yields a critic-free, actor-only least-squares update computed directly from input-state data. Under a veriﬁable projected rank condition, the resulting data equation is equivalent to the policy-space residual equation, and each update coincides with the Kleinman iteration. Thus, the stabilizing and convergence properties of Kleinman iteration are retained without a critic regression. We further show that the conventional off-policy full-rank condition decomposes into an endpoint critic rank condition and a projected actor rank condition. The proposed method removes the rank requirement needed for critic identiﬁcation while retaining the one needed for policy improvement. The repeated least-squares dimension is reduced from n (n + 1)/2 + mn to mn. Finally, comparative simulations validate the effectiveness of the proposed algorithm.  \nIndex Terms—Policy iteration, linear quadratic regulation, data-driven control, critic-free reinforcement learning, endpoint null-space projection.  \nI. INTRODUCTION  \nCONTINUOUS-time linear quadratic regulation is a fun  \ndamental problem in optimal control [1] . When the system matrices are known, the optimal feedback gain can be obtained by solving the algebraic Riccati equation (ARE) . In many practical applications, however, accurate knowledge of the system matrices is unavailable, which has motivated extensive research on data-driven optimal control [2], reinforcement learning (RL) [3], and adaptive dynamic programming [4] .  \nFor this problem, data-driven policy iteration (PI) provides a natural way to implement Riccati-based optimal control from data [5]–[7] . Starting from an admissible feedback gain, PI alternates between policy evaluation and policy improvement, and its model-based form is closely related to the classical Kleinman iteration. In the data-driven setting, Vrabie et al. [8] introduced an on-policy PI scheme based on integral RL, where the feedback gain is updated from closed-loop inputstate data rather than by directly solving the ARE. This approach reduces the reliance on a complete system model,  \nJiacheng Wu, Yang Zhu, and Hongye Su are with State Key Laboratory of Industrial Control Technology, Institute of Cyber-Systems and Control, Zhejiang University, Hangzhou 310027, China (e-mail: ji[achengwu@zju.edu.cn](achengwu@zju.edu.cn); [zhuyang88@zju.edu.cn](zhuyang88@zju.edu.cn); [hysu@iipc.zju.edu.cn](hysu@iipc.zju.edu.cn);).  \n* Corresponding author: Yang Zhu.  \nbut its policy-improvement step still requires partial knowledge of the system dynamics. To further reduce model dependence, Jiang et al. [9] developed an off-policy PI framework in which integral regression equations are constructed from trajectories generated by a behavior input, so that the unknown system matrices need not be explicitly identiﬁed. This framework has motivated a broad range of extensions and applications, including inverse optimal control [10], permanent magnet synchronous motor [11], singularly perturbed systems [12], and networked [13] or multi-agent systems [14], [15] . Nevertheless, the resulting regression treats the value matrix of the current policy and the improved feedback gain as simultaneous unknowns. As a result, each policy update still requires a joint critic–acto","cbCainhaQPkJkE76","https://ap.wps.com/l/cbCainhaQPkJkE76","pdf",361141,3,1,"English","en",105,"# Abstract and Index Terms\n# Introduction\n## Problem Background and Motivation\n## Data-Driven Policy Iteration and Critic–Actor Regression\n## Key Question: Is Critic Identification Necessary?\n## Related Policy-Space and Model-Based Results\n## Gap in Existing Off-Policy Integral Regression","[{\"question\":\"What theoretical results ensure the method’s stability and equivalence to Kleinman iteration?\",\"answer\":\"With a verifiable projected rank condition, the resulting data equation is equivalent to the policy-space residual equation, so each update coincides with the Kleinman iteration, retaining its stabilizing and convergence properties.\"}]",1784200710,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":77,"head_meta":79,"extra_data":81,"updated_unix":27},"data-driven-critic-free-policy-iteration-for-continuous-time-linear-quadratic-regulation","",{"@graph":35,"@context":76},[36,52,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,49],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":20},"https://docshare.wps.com/document/research-report/",{"item":50,"name":13,"@type":42,"position":51},"https://docshare.wps.com/document/data-driven-critic-free-policy-iteration-for-continuous-time-linear-quadratic-regulation/85060/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":23,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70],{"name":71,"@type":72,"acceptedAnswer":73},"What theoretical results ensure the method’s stability and equivalence to Kleinman iteration?","Question",{"text":74,"@type":75},"With a verifiable projected rank condition, the resulting data equation is equivalent to the policy-space residual equation, so each update coincides with the Kleinman iteration, retaining its stabilizing and convergence properties.","Answer","https://schema.org",{"og:url":50,"og:type":78,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":80,"canonical":50},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":83},[84,88,92,96,101,106,111,114,118,121,125],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":85,"show_sort_weight":86,"slug":87},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":89,"show_sort_weight":90,"slug":91},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Exam",70,"exam",{"id":97,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},5,"Comic",60,"comic",{"id":102,"doc_module":4,"doc_module_name":45,"category_name":103,"show_sort_weight":104,"slug":105},6,"Technology",50,"technology",{"id":107,"doc_module":4,"doc_module_name":45,"category_name":108,"show_sort_weight":109,"slug":110},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":112,"slug":113},30,"research-report",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":28,"slug":117},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":119,"show_sort_weight":28,"slug":120},"World Cup","world-cup",{"id":122,"doc_module":4,"doc_module_name":45,"category_name":123,"show_sort_weight":122,"slug":124},10,"Lifestyle","lifestyle",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":97,"slug":128},19,"General","general"]