[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83517-en":3,"doc-seo-83517-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83517,549758146520,"Patrick","https://ap-avatar.wpscdn.com/avatar/80002397d8c0411e94?_k=1775819394049821470",8,"Research & Report","A Data-Enabled Primal-Dual Approach for Policy Learning with SDP Formulations","This paper develops a data-enabled primal-dual framework for learning optimal control policies for unknown linear discrete-time systems from online data. The data-dependent control synthesis problem is formulated as a time-varying semidefinite program (SDP) whose coefficients are recursively updated using closed-loop measurements. Online policy updates avoid full SDP re-solves by using lightweight primal-dual iterations: a linear solve plus projection onto the positive semidefinite cone. Results quantify coupling via the Sim-to-Real Gap and Difference-of-Signal and establish local tracking and global convergence bounds, validated on LQR, H∞, and safety-critical exploration.","arXiv :2607 .00644v1 [ ee ss . SY] 1 Jul 2026  \nA Data-Enabled Primal-Dual Approach for Policy Learning with  \nSDP Formulations  \nHan Wang Feiran Zhao Florian D¨orﬂer  \nAutomatic Control Laboratory (IfA), ETH Z¨urich  \n{hanwang1,zhaofe,[dorfler](dorfler}@ethz. ch)[}](dorfler}@ethz. ch)[@ethz. ch](dorfler}@ethz. ch)  \nAbstract  \nThis paper develops a data-enabled primal-dual framework for learning optimal control policies for unknown linear discrete-time systems from online data. The proposed approach views the data-dependent control synthesis problem as a time-varying semideﬁnite program (SDP) whose coeﬃcients are recursively updated from online closed-loop measurements. Instead of repeatedly solving a full SDP as new data arrive, the policy is updated online through lightweight primaldual iterations, each consisting of a linear equation solve and a projection onto the positive semideﬁnite cone. The framework applies to both direct and indirect data-driven formulationsand covers a broad class of control objectives, including LQR, H∞ control, and safety-critical control. To characterize the coupling between online optimization and closed-loop data generation, we introduce two data-dependent quantities: the Sim-to-Real Gap, which measures themismatch between noisy and noiseless data-induced SDPs, and the Diﬀerence-of-Signal, which measures the temporal variation of the SDP coeﬃcients. Under persistency of excitation, suitable SDP regularity conditions, and suﬃciently slow data variation, we establish a local linear tracking result up to residual terms governed by the latter two quantities. A global ergodic convergence bound is also derived for arbitrary initialization. Numerical examples on LQR, H∞ control, and safe exploration demonstrate that the proposed method can eﬃciently improve control performance from online data while accommodating SDP constraints beyond the well-explored LQR policy-gradient formulations.  \nKeywords: Data-driven control, adaptive control.  \n1 Introduction  \nData-driven control aims to synthesize controllers from measured trajectories, thereby reducing the dependence on explicit ﬁrst-principles modeling. This paradigm is particularly attractive for modern control applications, where high-dimensional dynamics, uncertain environments, and repeated operation make it diﬃcult to obtain an accurate model beforehand, while closed-loop data are continuously generated during deployment. A central challenge is therefore not only how to design a controller from a ﬁxed batch of data, but also how to exploit newly collected data to recursively improve and adapt the learned control policy online.  \n1.1 Literature review  \nA key theoretical foundation of data-driven control is that measured trajectories can serve as an implicit representation of unknown dynamics, as formalized by Willems’ fundamental lemma under persistency of excitation [1] . This viewpoint has supported both predictive-control formulations such as DeePC [2–4] and data-dependent linear matrix inequalities methods for stabilization, LQR,  \nand robust control [5–7], with related informativity results clarifying when a dataset is suﬃcient to certify control properties [8] . These methods are typically used in a batch manner: a controller is synthesized from a ﬁxed dataset, while data collected during subsequent operation are not directly used to reﬁne the policy. This motivates a recursive learning viewpoint, where the policy is updated as new closed-loop data arrive, without repeatedly solving the full data-dependent synthesis problem from scratch.  \nRecent online policy-learning methods partially address this issue by combining closed-loop data with recursive controller updates, especially for LQR problems [9–12] . These methods provide an appealing mechanism for improving optimal control policies from online data and have also been validated in robotic applications [13] . Their algorithmic structure, however, is often tailored to LQR-type objectives ","cbCaikX9XLm1N0Lu","https://ap.wps.com/l/cbCaikX9XLm1N0Lu","pdf",729961,4,1,35,"English","en",105,"# Introduction\n## Literature review","[{\"question\":\"What problem does the paper address?\",\"answer\":\"It addresses learning optimal control policies for unknown linear discrete-time systems using online closed-loop data, without relying on accurate first-principles models.\"},{\"question\":\"How is the online learning problem formulated?\",\"answer\":\"The method treats data-dependent control synthesis as a time-varying semidefinite program (SDP) whose coefficients are recursively updated from ongoing measurements.\"},{\"question\":\"How does the proposed algorithm update the policy online?\",\"answer\":\"Instead of solving a full SDP each time new data arrive, it performs lightweight primal-dual iterations, combining a linear equation solve with a projection onto the positive semidefinite cone.\"}]",1784188568,88,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"a-data-enabled-primal-dual-approach-for-policy-learning-with-sdp-formulations","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/a-data-enabled-primal-dual-approach-for-policy-learning-with-sdp-formulations/83517/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address?","Question",{"text":75,"@type":76},"It addresses learning optimal control policies for unknown linear discrete-time systems using online closed-loop data, without relying on accurate first-principles models.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is the online learning problem formulated?",{"text":80,"@type":76},"The method treats data-dependent control synthesis as a time-varying semidefinite program (SDP) whose coefficients are recursively updated from ongoing measurements.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the proposed algorithm update the policy online?",{"text":84,"@type":76},"Instead of solving a full SDP each time new data arrive, it performs lightweight primal-dual iterations, combining a linear equation solve with a projection onto the positive semidefinite cone.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]