[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83364-en":3,"doc-seo-83364-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83364,1099514068365,"Aurelia","https://ap-avatar.wpscdn.com/avatar/10000253d8d9f28188e?_k=1776742907772140068",8,"Research & Report","Spectral Analysis of Dueling Q-Learning","Q-learning is a reinforcement learning method for solving discounted Markov decision processes with unknown transition kernels. Deep Q-networks extend it by approximating the Q-function with deep neural networks, while dueling architectures split the approximation into a value stream and an advantage stream to improve learning efficiency. Theoretical understanding remains incomplete, especially for pure tabular updates. This paper provides an exact switching linear system representation for deterministic dueling Q-learning and proves finite-time expected error bounds for the sampled stochastic recursion, clarifying how value and advantage updates act as distinct gains on Q’s action-common and action-differential components.","arXiv :2607 .08340v 1 [ cs .LG] 9 Jul 2026  \nSpectral Analysis of Dueling Q-Learning  \nDonghwan Lee  \nDepartment of Electrical Engineering, Korea Advanced Institute of Science and Technology (KAIST)  \nDaejeon 34141, South Korea (email: [donghwan@kaist.ac.kr](donghwan@kaist.ac.kr))  \nAbstract  \nQ-learning is a fundamental algorithm in reinforcement learning (RL) for solving discounted Markov decision processes (MDPs) when the transition kernel is unknown. The deep Q-network (DQN) extends Q-learning by using a deep neural network for Q-function approximation, which makes Q-learning applicable to more practical high-dimensional problems. Dueling Q-learning decomposes the Q-function into a value function and an advantage function and learns the two components jointly, which can improve learning eﬃciency. However, the theoretical understanding of dueling Q-learning is still limited. Recent work has initiated an analysis of tabular dueling Q-learning, but existing guarantees focus on a regularized formulation and leave the pure tabular update less completely understood. This paper strengthens that line of analysis by adding a direct interpretation of the centered tabular decomposition and by establishing convergence guarantees for the unregularized, unprojected constant step-size recursion. In particular, we derive an exact switching linear system representation for deterministic dueling Q-learning anda ﬁnite-time error bound in expectation for the sampled stochastic version. The analysis clariﬁes how the value and advantage updates act as diﬀerent gains on the action-common (value function) and action-diﬀerential (advantage function) components of the Q-function.  \n1 Introduction  \nQ-learning [7] is a foundational algorithm in reinforcement learning (RL) [10] for solving discounted Markov decision processes (MDPs) [8] with unknown transition kernels. The deep Q-network (DQN) [14] extends Q-learning by using a deep neural network for Q-function approximation, which makes value-based RL applicable to high-dimensional problems in which a tabular representation is not practical. The dueling network architecture [15] further modiﬁes DQN by separating the Q-function approximation into a value stream and an advantage stream. This decomposition can improve learning eﬃciency because the value component can share state-wise information across actions while the advantage component captures action-dependent deviations. Despite its empirical usefulness, the theoretical understanding of dueling Q-learning is still much less complete than that of standard tabular Q-learning.  \nA recent study of action-value temporal-diﬀerence methods [13] that learn state values formalized tabular versions of dueling methods and introduced regularized dueling Q-learning. That work clariﬁed important aspects of dueling Q-learning (also called AV-learning) and provided a theoretical solution analysis for a regularized formulation of dueling Q-learning. However, the pure tabular dueling Q-learning recursion, without a regularization term or projection step, still leaves room for a more direct convergence analysis. The present paper addresses this gap by analyzing the unregularized dueling Q-learning recursion with constant step-sizes and by deriving a ﬁnite-time error bound in expectation for its sampled stochastic version.  \nThe analysis begins by interpreting dueling Q-learning through an orthogonal common and diﬀerential decomposition of the tabular Q-function. For a tabular Q-function Q, deﬁne the statewise mean  \n¯VQ (s) := |~~ ~~1A| X Q(s, b),  \nb∈A  \nand identify this state vector with its action-repeated lift  \n(ΠQ)(s, a) := ¯VQ (s), s ∈ S, a ∈ A.  \nThus ΠQ is the projection of Q onto the subspace of vectors that are constant across actions in each state. The action-diﬀerential component is the complementary centered projection  \nA (s, a) = Q(s, a) − (ΠQ)(s, a) = Q(s, a) − |~~ ~~1A| X Q(s, b) .  \nb∈A  \nTherefore, Q is written as the sum of two components, de","cbCaijnXEbsiBXV3","https://ap.wps.com/l/cbCaijnXEbsiBXV3","pdf",437537,3,1,43,"English","en",105,"# Introduction\n## Dueling Q-learning motivation and gap\n## Orthogonal common–differential decomposition\n## Algorithm variants and equivalence\n## Deterministic vs sampled stochastic analysis\n## Switching linear system framework","[{\"question\":\"What decomposition does the paper use to analyze dueling Q-learning?\",\"answer\":\"It decomposes the tabular Q-function into an action-common (value) component and an action-differential (advantage) component using orthogonal centered projections.\"},{\"question\":\"What main results does the paper establish about convergence?\",\"answer\":\"It derives a switching linear system representation and proves finite-time error bounds in expectation for the sampled stochastic dueling Q-learning recursion, focusing on unregularized and unprojected constant step-size updates.\"},{\"question\":\"How do the value and advantage updates differ in the analysis?\",\"answer\":\"The paper shows they act with different effective gains on two orthogonal subspaces: the action-common component and the action-differential component, clarifying the learning dynamics behind the dueling architecture.\"}]",1784187009,108,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"spectral-analysis-of-dueling-q-learning","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/spectral-analysis-of-dueling-q-learning/83364/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What decomposition does the paper use to analyze dueling Q-learning?","Question",{"text":75,"@type":76},"It decomposes the tabular Q-function into an action-common (value) component and an action-differential (advantage) component using orthogonal centered projections.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What main results does the paper establish about convergence?",{"text":80,"@type":76},"It derives a switching linear system representation and proves finite-time error bounds in expectation for the sampled stochastic dueling Q-learning recursion, focusing on unregularized and unprojected constant step-size updates.",{"name":82,"@type":73,"acceptedAnswer":83},"How do the value and advantage updates differ in the analysis?",{"text":84,"@type":76},"The paper shows they act with different effective gains on two orthogonal subspaces: the action-common component and the action-differential component, clarifying the learning dynamics behind the dueling architecture.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]