[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85354-en":3,"doc-seo-85354-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85354,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","Active Offline-to-Online Reinforcement Learning","Active offline-to-online reinforcement learning (O2O-RL) trains effective policies from large offline datasets, then improves them with limited online interaction in settings where exploration is costly or risky. The document targets active policy selection for fine-tuning under a constrained interaction budget, addressing sensitivity to algorithm and hyperparameters in standard pipelines. It formulates a trade-off between interaction for policy evaluation and for fine-tuning, and proposes an active method using upper-confidence bounds from locally linear performance forecasts.","arXiv :2607 . 1 1720v 1 [ cs .LG] 13 Jul 2026  \nActive Offline-to-Online Reinforcement Learning  \nALPER KAMIL BOZKURT* , Virginia Commonwealth University, USA SHANGTONG ZHANG, University of Virginia, USA  \nYUICHI MOTAI, Virginia Commonwealth University, USA  \nBackground: Offline reinforcement learning (RL) enables effective policies to be trained from large, previously collected datasets and subsequently improved through limited online interaction. This offline-to-online RL (O2O-RL) paradigm is particularly promising in nonstationary domains where interaction is costly or potentially hazardous. Standard O2O-RL pipelines train multiple candidate policies offline, evaluate them using off-policy or online evaluation, and then deploy and fine-tune the policy with the highest estimated value. However, as in offline pretraining, fine-tuning performance is highly sensitive to the choice of algorithm and hyperparameters, making it risky to commit to a single policy.  \nObjectives: We study active policy selection for fine-tuning under a limited interaction budget in O2O-RL settings. To our knowledge, this is the first work to address this problem.  \nMethods: We formulate the problem by identifying a fundamental trade-off between allocating online interactions to policy evaluation, which helps identify high-performing policies, and allocating them to fine-tuning, which improves policy performance. We then propose an approach that balances this trade-off by actively selecting policies for fine-tuning based on upper-confidence bounds on their future performance. These bounds are derived from locally linear performance forecasts fitted to observations obtained through online evaluation.  \nResults: Across a diverse range of experiments, the proposed approach consistently outperforms existing O2O-RL baselines. Conclusions: Actively selecting and fine-tuning policies uses limited online interaction budgets more effectively than either committing to a single policy or dividing the budget equally among all policies. Our framework also advances offline RL toward practical deployment in real-world systems where online interaction is costly or risky.  \n1 Introduction  \nReinforcement learning (RL) is becoming a key ingredient in the autonomy of modern robotic systems operating in unstructured, dynamic environments (Singh et al. 2022) . By learning to make decisions and derive control actions directly from onboard sensing and perception, RL can substantially reduce human workload and the likelihood of human error (Zhang et al. 2022), thereby facilitating widespread real-world deployment. Deep RL has proven effective in synthesizing control policies for high-dimensional, nonlinear physical systems for which manual controller design is infeasible, leading to numerous successful applications (Tang et al. 2025). Despite the flexibility and power of this framework, standard RL methods typically require extensive direct interaction with the physical environment for exploration (Ladosz et al. 2022) . They are therefore impractical for training policies from scratch when such interactions are costly, risky, or time-consuming (Dulac-Arnold et al. 2021) .  \nOffline RL (Levine et al. 2020 ; Prudencio et al. 2023) has emerged as an alternative to online RL, enabling policies to be trained from large, previously collected datasets. These datasets are usually collected under safe, controlled conditions (G. Zhou et al. 2023), often, though not exclusively, by human operators. A central challenge in offline RL is that the performance of a pretrained policy becomes unpredictable as its behavior diverges from that of the policy used to collect the dataset (Kostrikov et al. 2021) . Due to this distributional shift, a policy pretrained via offline RL can perform arbitrarily poorly in the real environment (Qin et al. 2022) . To partially mitigate this issue, offline policy selection, usually performed via off-policy evaluation (OPE) (Paine et al. 2020 ; Uehara et al. 20","cbCainexRfIYhC9g","https://ap.wps.com/l/cbCainexRfIYhC9g","pdf",578516,1,18,"English","en",105,"# Introduction\n## Offline RL and distributional shift\n## Offline-to-online RL (O2O-RL) paradigm\n## Policy selection and evaluation challenges\n# Objectives\n# Methods\n# Results\n# Conclusions","[{\"question\":\"What is offline-to-online reinforcement learning (O2O-RL)?\",\"answer\":\"O2O-RL trains policies from large offline datasets and then improves them using a small amount of online interaction. It is especially useful when environments are nonstationary and online interaction is costly or hazardous.\"},{\"question\":\"Why is fine-tuning in O2O-RL risky?\",\"answer\":\"Fine-tuning performance is highly sensitive to algorithm choices and hyperparameters. As a result, committing to a single policy can lead to unreliable outcomes.\"},{\"question\":\"How does the proposed method balance interaction usage in O2O-RL?\",\"answer\":\"It identifies a trade-off between allocating interactions to policy evaluation and allocating them to fine-tuning. The method actively selects policies for fine-tuning using upper-confidence bounds derived from locally linear performance forecasts fitted to online evaluation observations.\"}]",1784202743,45,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"active-offline-to-online-reinforcement-learning","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/active-offline-to-online-reinforcement-learning/85354/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is offline-to-online reinforcement learning (O2O-RL)?","Question",{"text":75,"@type":76},"O2O-RL trains policies from large offline datasets and then improves them using a small amount of online interaction. It is especially useful when environments are nonstationary and online interaction is costly or hazardous.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why is fine-tuning in O2O-RL risky?",{"text":80,"@type":76},"Fine-tuning performance is highly sensitive to algorithm choices and hyperparameters. As a result, committing to a single policy can lead to unreliable outcomes.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the proposed method balance interaction usage in O2O-RL?",{"text":84,"@type":76},"It identifies a trade-off between allocating interactions to policy evaluation and allocating them to fine-tuning. The method actively selects policies for fine-tuning using upper-confidence bounds derived from locally linear performance forecasts fitted to online evaluation observations.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]