[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85612-en":3,"doc-seo-85612-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85612,3848291630094,"Emma Wilson","https://eur-avatar.wpscdn.com/davatar_085a072bc5b1113ac321206ff7593b45",8,"Research & Report","Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification","We study KL-regularized contextual bandits and episodic reinforcement learning under general function approximation with model misspecification. Existing analyses depend on realizability and therefore break when misspecified models introduce approximation bias. The work formulates KL-aligned misspecification conditions for bandits and episodic RL and analyzes optimistic, regression-based methods with Gibbs policy updates. Results provide high-probability KL-regret guarantees with explicit misspecification terms, recovering the realizable KL-regularized setting as a special case.","arXiv :2606 .06053v2 [ cs .LG] 12 Jul 2026  \nOnline KL-Regularized Reinforcement Learning with Function Approximation under Misspecification  \nHaoyang Hong∗ , Zichen Wang∗ , Quanquan Gu, Huazheng Wang  \nKeywords: KL regularization, contextual bandits, reinforcement learning, model misspecification  \nSummary  \nWe study KL-regularized contextual bandits and episodic reinforcement learning with reference policies under general function approximation and model misspecification, motivated by KL-constrained policy optimization and RLHF-style post-training (Schulman et al., 2017 ; Ouyang et al., 2022 ; Rafailov et al., 2023 ; Zhao et al., 2025) . The key difficulty is that KL regularization changes the benchmark, so standard misspecification conditions do not transfer directly. We introduce KL misspecification conditions and analyze optimistic regression-based algorithms with Gibbs policy updates. Our main results are a high-probability KL-regret guarantee for contextual bandits and a modular leading-order high-probability regret theorem for episodic RL under explicit confidence and uncertainty conditions, both with explicit dependence on misspecification and statistical complexity, recovering the corresponding realizable guarantees of Zhao et al. (2025) as special cases.  \nContribution(s)  \n1. We introduce KL-aligned misspecification conditions for KL-regularized contextual banditsand episodic reinforcement learning under general function approximation.  \nContext: The conditions are motivated by the Gibbs form of the KL-regularized objective and complement prior KL-regularized analyses that focus on realizable or near-realizable settings (Zhao et al., 2025) . Our contextual-bandit condition is a pointwise misspecification control over all state-action pairs, paralleling classical contextual-bandit misspecification (Foster et al., 2021b ; Takemura et al., 2021) . Our RL condition extends this pointwise control to stagewise KL Bellman-backup targets.  \n2. We study oracle-based optimistic algorithms for KL-regularized contextual bandits and episodic reinforcement learning, combining regression-based estimation, uncertainty bonuses, and Gibbs policy updates relative to a reference policy.  \nContext: The algorithms follow the optimistic reward-estimation template and are compatible with the KL-regularized objective. Our primary guarantees are for the knownmisspecification setting, where the misspecification levels enter the bonuses as inputs.  \n3. We establish high-probability KL-regret guarantees with explicit misspecification terms and complexity-sensitive dependence, including a direct bandit theorem and a high-probability regret theorem for episodic RL stated under assumed confidence and uncertainty conditions. Context: The guarantees are stated against the KL-optimal Gibbs comparator and use localized eluder-dimension-style complexity terms. For RL, the theorem is modular and is under explicit confidence and uncertainty conditions.  \nOnline KL-Regularized Reinforcement Learning with Function Approximation under Misspecification  \nHaoyang Hong 1 ∗ , Zichen Wang2∗, Quanquan Gu3 , Huazheng Wang 1 [honghao@oregonstate.edu](honghao@oregonstate.edu) , [zichenw6@illinois.edu](zichenw6@illinois.edu) , [qgu@cs.ucla.edu](qgu@cs.ucla.edu) , [huazheng.wang@oregonstate.edu](huazheng.wang@oregonstate.edu)  \n1 School of Electrical Engineering and Computer Science, Oregon State University  \n2Department of Electrical and Computer Engineering and Coordinated Science Laboratory, University of Illinois Urbana-Champaign  \n3Department of Computer Science, University of California, Los Angeles  \nAbstract  \nWe study KL-regularized contextual bandits and episodic reinforcement learning (RL) under general function approximation with model misspecification. Existing guarantees rely on realizability and therefore do not extend to misspecified models, where classical regret bounds may fail. This work introduces KL misspecification formulations for contextual ba","cbCaigKNUl1XeWG1","https://ap.wps.com/l/cbCaigKNUl1XeWG1","pdf",479810,2,1,32,"English","en",105,"# Summary\n# Contribution(s)\n# Abstract\n# 1 Introduction\n## Motivation in RLHF and KL-penalized policy optimization\n## What misspecification means under function approximation\n## Related work and analytical foundation","[{\"question\":\"What problem does the paper address in KL-regularized online RL?\",\"answer\":\"It studies KL-regularized contextual bandits and episodic reinforcement learning under general function approximation when the model class is misspecified. The goal is to obtain regret guarantees despite approximation bias that breaks realizability-based analyses.\"},{\"question\":\"How does KL regularization affect misspecification analysis?\",\"answer\":\"KL regularization changes the benchmark via a KL-constrained objective, so standard misspecification conditions do not carry over directly. The paper introduces KL-aligned misspecification conditions tailored to the Gibbs-form KL-regularized objective.\"},{\"question\":\"What kind of algorithms and guarantees are proposed?\",\"answer\":\"The paper analyzes optimistic, oracle-based regression methods that combine regression-based estimation, uncertainty bonuses, and Gibbs policy updates relative to a reference policy. It establishes high-probability KL-regret guarantees with explicit dependence on misspecification and complexity-sensitive terms.\"}]",1784204920,81,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"online-kl-regularized-reinforcement-learning-with-function-approximation-under-misspecification","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/online-kl-regularized-reinforcement-learning-with-function-approximation-under-misspecification/85612/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in KL-regularized online RL?","Question",{"text":75,"@type":76},"It studies KL-regularized contextual bandits and episodic reinforcement learning under general function approximation when the model class is misspecified. The goal is to obtain regret guarantees despite approximation bias that breaks realizability-based analyses.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does KL regularization affect misspecification analysis?",{"text":80,"@type":76},"KL regularization changes the benchmark via a KL-constrained objective, so standard misspecification conditions do not carry over directly. The paper introduces KL-aligned misspecification conditions tailored to the Gibbs-form KL-regularized objective.",{"name":82,"@type":73,"acceptedAnswer":83},"What kind of algorithms and guarantees are proposed?",{"text":84,"@type":76},"The paper analyzes optimistic, oracle-based regression methods that combine regression-based estimation, uncertainty bonuses, and Gibbs policy updates relative to a reference policy. It establishes high-probability KL-regret guarantees with explicit dependence on misspecification and complexity-sensitive terms.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]