[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-127519-en":3,"doc-seo-127519-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},127519,13056712833777,"Logic","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Quantum Natural Policy Gradients - Towards Sample-Efficient Reinforcement Learning","Reinforcement learning enables intelligent behavior through trial-and-error interaction, but this process often requires large numbers of environment samples. Variational quantum circuits can reduce training cost by acting as function approximators on current noisy quantum hardware. The proposed quantum natural policy gradient (QNPG) is a second-order policy-gradient method that efficiently approximates the quantum Fisher information matrix. Experiments on contextual bandits show improved convergence speed and stability versus first-order training, reducing sample complexity, and the approach is validated on a 12-qubit hardware device.","Quantum Natural Policy Gradients: Towards Sample-Efﬁcient Reinforcement Learning  \nNico Meyer􀀃y , Daniel D. Scherer􀀃 , Axel Plinge􀀃 , Christopher Mutschler􀀃 , Michael J. Hartmanny  \n􀀃 Fraunhofer IIS, Fraunhofer Institute for Integrated Circuits IIS, N¨urnberg, Germany y Friedrich-Alexander University Erlangen-N¨urnberg (FAU), Department of Physics, Erlangen, Germany  \narXiv :2304 . 13571v1 [ quant-ph] 26 Apr 2023  \nAbstract—Reinforcement learning is a growing ﬁeld in AI with a lot of potential. Intelligent behavior is learned automatically through trial and error in interaction with the environment. However, this learning process is often costly. Using variational quantum circuits as function approximators can reduce this cost. In order to implement this, we propose the quantum natural policy gradient (QNPG) algorithm – a second-order gradient-based routine that takes advantage of an efﬁcient approximation of the quantum Fisher information matrix. We experimentally demonstrate that QNPG outperforms ﬁrst-order based training on Contextual Bandits environments regarding convergence speed and stability and thereby reduces the sample complexity. Furthermore, we provide evidence for the practical feasibility of our approach by training on a 12-qubit hardware device.  \nIndex Terms—reinforcement learning, variational quantum computing, policy gradient, natural gradient, contextual bandits  \nI. INTRODUCTION  \nOne critical technical factor in both classical and quantum reinforcement learning (RL) is the sample complexity, as interaction with the environment is potentially costly. Enhancing RL with variational quantum circuits (VQCs) as function approximators is a promising approach to reduce this cost utilizing the current noisy quantum hardware. Variational quantum algorithms are considered one of the prime applications for noisy intermediate-scale quantum (NISQ) computing.  \nThe concept can be leveraged as a platform for quantum machine learning (QML) [1], which provides a provable quantum advantage for speciﬁc problems [2], [3] . Concrete realizations typically combine a VQC with a classical training routine. This approach is believed to have some robustness to the (currently) inevitable hardware noise [4],[5] . VQC parameter updates can be computed using ﬁrst-order gradients [6] .  \nQuantum reinforcement learning (QRL) [7] aims at enhancing classical reinforcement learning [8] with quantum computing. NISQ-compatible instances of QRL employ VQC-based  \nThe research was supported by the Bavarian Ministry of Economic Affairs, Regional Development and Energy with funds from the Hightech Agenda Bayern via the project BayQS and by the Bavarian Ministry for Economic Affairs, Infrastructure, Transport and Technology through the Center for AnalyticsData-Applications (ADA-Center) within the framework of “BAYERN DIGITAL II”.  \nM. Hartmann acknowledges support by the European Union's Horizon 2020 research and innovation programme under grant agreement No 828826 “Quromorphic” and the Munich Quantum Valley, which is supported by the Bavarian state government with funds from the Hightech Agenda Bayern Plus. Corresponding author: [nico.meyer@iis.fraunhofer.de](nico.meyer@iis.fraunhofer.de)  \nFig. 1. Proposed method: The update with the ﬁrst-order gradient r 􀀒􀀙 􀀒 is extended with a second-order term g (􀀒) . This deﬁnes a (quantum) natural gradient approach, which aims for training in a partially undistorted neighborhood of the parameter space – improving convergence behavior.  \nfunction approximators for quantum Q-learning [9] and quantum policy gradient (QPG) [9] approaches.  \nA concern for both QML and QRL is the trainability of the VQC, and the associated sample complexity [10], i.e., the required interactions with the environment. To enable a more targeted training procedure [11],[12], one can include secondorder terms in the parameter update – at the expense of circuit evaluation overhead.  \nContribution. We propose a second-order extension 1 ","cbCaibFfK5WUOZHq","https://ap.wps.com/l/cbCaibFfK5WUOZHq","pdf",994222,1,7,"English","en",105,"# Introduction\n# Method\n## Markov Decision Process formulation\n## Proposed QNPG approach\n# Related Work\n# Empirical Results\n## Contextual bandits proof-of-concept\n## 12-qubit hardware evaluation","[{\"question\":\"Why is sample complexity a key challenge in reinforcement learning for this work?\",\"answer\":\"Interaction with the environment can be costly, making the learning process sample-intensive. Reducing sample complexity is the central motivation.\"},{\"question\":\"What is the quantum natural policy gradient (QNPG) algorithm?\",\"answer\":\"QNPG is a second-order gradient-based routine built from policy gradients, using an efficient approximation of the quantum Fisher information matrix.\"},{\"question\":\"How does QNPG perform compared with first-order training and on what setups?\",\"answer\":\"On contextual bandits, QNPG improves convergence speed and stability over first-order training, and it also performs well on a 12-qubit hardware device.\"}]","Quantum Natural Policy Gradients - Towards Sample-Efficient Reinforcement Learning | PDF",1785939708,18,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"quantum-natural-policy-gradients-towards-sample-efficient-reinforcement-learning","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/quantum-natural-policy-gradients-towards-sample-efficient-reinforcement-learning/127519/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-23","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why is sample complexity a key challenge in reinforcement learning for this work?","Question",{"text":76,"@type":77},"Interaction with the environment can be costly, making the learning process sample-intensive. Reducing sample complexity is the central motivation.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What is the quantum natural policy gradient (QNPG) algorithm?",{"text":81,"@type":77},"QNPG is a second-order gradient-based routine built from policy gradients, using an efficient approximation of the quantum Fisher information matrix.",{"name":83,"@type":74,"acceptedAnswer":84},"How does QNPG perform compared with first-order training and on what setups?",{"text":85,"@type":77},"On contextual bandits, QNPG improves convergence speed and stability over first-order training, and it also performs well on a 12-qubit hardware device.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]