[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82884-en":3,"doc-seo-82884-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82884,8796095462418,"Noah","https://ap-avatar.wpscdn.com/avatar/80000253c1241d02b47?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778826106357471780",8,"Research & Report","Non-Convex Sparse Reinforcement Learning via Non-Monotone Inclusions","Non-Convex Sparse Reinforcement Learning via Non-Monotone Inclusions presents two contributions: sparsity-driven feature selection for reinforcement learning (RL) and new theory for non-monotone inclusions. For RL, it reduces estimation bias in classical regularization by augmenting least-squares temporal-difference (LSTD) policy evaluation with a non-convex projected minimax concave (PMC) penalty. Because the PMC penalty is weakly convex, the fixed-point problem becomes non-monotone, leading to FRBS convergence results. Experiments on benchmark data show strong gains in noisy, high-dimensional feature regimes.","Non-Convex Sparse Reinforcement Learning via  \nNon-Monotone Inclusions  \nKyohei Suzuki, Member, IEEE, and Konstantinos Slavakis, Senior Member, IEEE  \narXiv :2607 .04990v2 [ cs .LG] 7 Jul 2026  \nAbstract—This work delivers two key contributions: one to efficient feature selection in reinforcement learning (RL), the other to the theory of non-monotone inclusions. On the RL side, the estimation bias inherent in conventional regularization schemes is addressed by augmenting classical least-squares temporal-difference (LSTD) policy evaluation with the sparsityinducing, non-convex projected minimax concave (PMC) penalty. Because the PMC penalty is weakly convex, the resulting fixedpoint problem is no longer monotone; instead, it falls under abroader class of non-monotone inclusions involving the sum of a monotone Lipschitz operator and a hypomonotone operator. On the theory side, novel convergence conditions are developed for the forward-reflected-backward splitting (FRBS) method applied to this broader class of non-monotone inclusion problems. Under mild conditions, Lyapunov stability and the existence of a limit point of the sequence of FRBS iterates are established; alternatively, under the weak Minty variational inequality assumption, exact convergence is guaranteed. Numerical tests on benchmark datasets show that the proposed FRBS iterates, applied to thenon-convexly regularized LSTD problem, substantially outperform state-of-the-art feature-selection methods, especially when many noisy features are present.  \nIndex Terms—Reinforcement learning, feature selection, sparse modeling, non-convex, non-monotone.  \nI. INTRODUCTION  \nREINFORCEMENT learning (RL) plays a central role  \nin contemporary machine learning, signal processing, and control theory [1]–[3] . Its central goal is for an agent to learn, through interaction with an environment typically modeled as a Markov decision process (MDP), an optimal policy that minimizes a long-term loss captured by the Qfunction. Yet in many real-world settings, such as robotics, educational agents, and healthcare, repeated online interaction with the environment is costly, and the collected data are often corrupted by erroneous measurements or outliers that can severely degrade the learned policy [4] . Robustness to such data imperfections is therefore essential, which motivates batch (offline) RL, where the agent learns from a fixed, precollected dataset of transitions rather than through additional online interaction.  \nIn most high-dimensional, real-world problems, explicitly representing the Q-function for all possible states and actions is impractical due to the “curse of dimensionality.”  \nK. Suzuki and K. Slavakis are with the Department of Information and Communications Engineering, Institute of Science Tokyo, Yokohama, 226- 8501, Japan. E-mails: [suzuki.k.439f@m.isct.ac.jp](suzuki.k.439f@m.isct.ac.jp), [slavakis@ict.eng.isct.ac.jp](slavakis@ict.eng.isct.ac.jp).  \nThis work was supported by the Grants-in-Aid for Scientific Research (KAKENHI) under Grant Number 25K24422 . This study was carried out using the TSUBAME4.0 supercomputer at Institute of Science Tokyo. An earlier version of this paper was presented at a conference [DOI: 10. 1109/ICASSP55912 .2026. 11463148] .  \nA common remedy is to approximate the Q-function using a functional (non-)parametric representation. This, however, introduces a fundamental trade-off between approximation accuracy and computational complexity: reducing the approximation error generally requires a large number of features in the model, which in turn increases computational demands. This manuscript focuses on “linear function” approximation—“linear” in some feature space—which, despite the success of deep RL, remains actively studied for its interpretability and its amenability to rigorous convergence analysis.  \nA. Feature selection in RL  \nFeature selection, achieved via a sparse representation over a large basis of functions, is an effective way","cbCaijwRenocy0Zs","https://ap.wps.com/l/cbCaijwRenocy0Zs","pdf",546152,1,17,"English","en",105,"# Abstract\n# Introduction\n## Feature selection in RL\n## Sparse regression","[{\"question\":\"What are the two main contributions of the work?\",\"answer\":\"The work provides (1) an RL method for efficient feature selection using a non-convex sparsity penalty, and (2) convergence theory for a forward-reflected-backward splitting (FRBS) method applied to non-monotone inclusions.\"},{\"question\":\"How does the proposed method address estimation bias in regularized RL?\",\"answer\":\"It augments least-squares temporal-difference (LSTD) policy evaluation with a projected minimax concave (PMC) penalty, which reduces the estimation bias caused by conventional regularization schemes.\"},{\"question\":\"What convergence guarantees are established for FRBS, and under what conditions?\",\"answer\":\"Under mild conditions it proves Lyapunov stability and existence of a limit point for the FRBS iterates; under a weak Minty variational inequality assumption it guarantees exact convergence.\"}]",1784183649,43,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"non-convex-sparse-reinforcement-learning-via-non-monotone-inclusions","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/non-convex-sparse-reinforcement-learning-via-non-monotone-inclusions/82884/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What are the two main contributions of the work?","Question",{"text":75,"@type":76},"The work provides (1) an RL method for efficient feature selection using a non-convex sparsity penalty, and (2) convergence theory for a forward-reflected-backward splitting (FRBS) method applied to non-monotone inclusions.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed method address estimation bias in regularized RL?",{"text":80,"@type":76},"It augments least-squares temporal-difference (LSTD) policy evaluation with a projected minimax concave (PMC) penalty, which reduces the estimation bias caused by conventional regularization schemes.",{"name":82,"@type":73,"acceptedAnswer":83},"What convergence guarantees are established for FRBS, and under what conditions?",{"text":84,"@type":76},"Under mild conditions it proves Lyapunov stability and existence of a limit point for the FRBS iterates; under a weak Minty variational inequality assumption it guarantees exact convergence.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]