[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86431-en":3,"doc-seo-86431-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86431,549758252649,"Ivy","https://ap-avatar.wpscdn.com/avatar/8000253669c5317157?_k=1778319167496531819",8,"Research & Report","RL Forgets! Towards Continual Policy Optimization","Continual post-training is a key paradigm for adapting vision-language models to evolving tasks, and reinforcement learning is often preferred over supervised fine-tuning based on claims of reduced forgetting. That claim lacks robust validation because existing results rely on outdated or narrow multimodal benchmarks. This work revisits the assumption using recent diverse reasoning tasks and introduces MRCL. Experiments show standard RL still triggers severe catastrophic forgetting. The failure is traced to an objective mismatch, and Continual Policy Optimization (CPO) mitigates drift without replay while preserving, sometimes improving, pretrained capabilities.","arXiv :2607 .04364v2 [ cs .LG] 13 Jul 2026  \nRL FORGETS! TOWARDS CONTINUAL POLICY OPTIMIZATION  \nMao-Lin Luo 1,2  Zhe-Xu Wang 1,2  Zi-Hao Zhou 1,2  Bo Ye 1,2,3 , Jian Zhao3,4 , Min-Ling Zhang 1,2 , Tong Wei 1,2†  \n1 School of Computer Science and Engineering, Southeast University, Nanjing 210096, China  \n2 Key Laboratory of Computer Network and Information Integration (Southeast University), Ministry of Education, China  \n3Zhongguancun Academy  \n4Zhongguancun Institute of Artificial Intelligence  \nABSTRACT  \nContinual post-training is becoming a central paradigm for adapting visionlanguage models to evolving tasks. Recent work has increasingly favored reinforcement learning over supervised fine-tuning, driven by the belief that reinforcement learning is inherently less prone to forgetting. However, the belief remains insufficiently validated, as existing evidence is largely drawn from outdated or homogeneous benchmarks. We revisit this assumption under recent and diverse multimodal reasoning tasks. To this end, we introduce MRCL, a Multimodal Reasoning Continual Learning benchmark. Experiments on MRCL show that standard reinforcement learning still suffers from severe catastrophic forgetting during continual post-training. We trace this failure to an objective mismatch:  \nthe KL regularization used in common policy optimization methods is evaluated on current-task data, whereas forgetting is caused by behavioral drift on priortask distributions. To address this problem, we propose Continual Policy Optimization (CPO), a replay-free framework grounded in a prior-task behavioral KL objective. CPO relaxes the intractable historical KL constraint into sparse parameter-movement regularization, limiting policy drift without storing old data.  \nExtensive experiments across multiple model scales show that CPO consistently reduces forgetting while preserving, and in some cases improving, pretrained model capabilities. On Qwen3-VL-8B, CPO reduces forgetting by 13.7% and improves pretrained capability by 7.0% . The implementation code is available at [https://github.com/MaolinLuo/CPO](https://github.com/MaolinLuo/CPO).  \n1 INTRODUCTION  \nVision-language models (VLMs) have demonstrated strong performance across a wide range of real-world tasks (Comanici et al., 2025; Achiam et al., 2023; Radford et al., 2021) . As VLMs are increasingly deployed in dynamic environments, they require continually adapt to new data and tasks. However, such a process is hindered by the fundamental challenge of catastrophic forgetting (French, 1999), where learning new tasks degrades previously acquired capabilities. Most prior work addresses this problem under supervised fine-tuning (SFT) (Gu et al., 2026a; Chen et al., 2025b; Kurniawan et al., 2024) . Recently, reinforcement learning (RL) has emerged as a promising alternative for continual post-training (Lai et al., 2025), with several studies reporting substantially less forgetting than SFT (Shenfeld et al., 2026; Zhang et al., 2026; Jin et al., 2025; Chen et al., 2025a) . These findings suggest an appealing possibility: is RL intrinsically resistant to forgetting in continual post-training?  \nWe show that the current evidence is incomplete. Existing evaluations have two blind spots that can make forgetting appear milder than it is. First, several benchmarks are old relative to the base models  \nbeing adapted. For example, Lai et al. (2025) train on datasets released between 2018 and 2023 while *Equal contribution.  \n†Corresponding author.  \nFigure 1: Plasticity–stability Pareto frontier during continual post-training. The y-axis reports the cumulative mean accuracy improvement (CAI) on current tasks, while the x-axis reports the cumulative mean accuracy delta (CAD) on previously learned tasks. Definitions of CAI and CAD are in Appendix B. Both RL and SFT exhibit catastrophic forgetting, whereas our method preserves prior knowledge more effectively while maintaining plasticity.  \nusing Qwen2.5-VL (Bai et al.","cbCaifnzrUUUouMI","https://ap.wps.com/l/cbCaifnzrUUUouMI","pdf",2774260,5,1,30,"English","en",105,"# Introduction\n## Continual post-training and catastrophic forgetting\n## Evidence gaps in existing RL evaluations\n## MRCL benchmark and experimental findings\n# Objective mismatch analysis\n## KL regularization evaluated on current-task data\n## Forgetting caused by behavioral drift on prior-task distributions\n# Continual Policy Optimization (CPO)\n## Replay-free prior-task behavioral KL objective\n## Sparse parameter-movement regularization","[{\"question\":\"What problem does the paper address in continual post-training for vision-language models?\",\"answer\":\"The paper addresses catastrophic forgetting, where learning new tasks during continual post-training degrades previously acquired capabilities.\"},{\"question\":\"Why do the authors argue that existing evidence for “RL forgets less” is incomplete?\",\"answer\":\"Existing evaluations can underestimate forgetting due to outdated/homogeneous benchmarks and narrow task distributions that do not reflect modern multimodal reasoning workloads.\"},{\"question\":\"How does Continual Policy Optimization (CPO) reduce forgetting differently from standard policy optimization?\",\"answer\":\"CPO replaces an intractable historical constraint with a replay-free framework based on a prior-task behavioral KL objective, limiting policy drift via sparse parameter-movement regularization rather than evaluating KL on current-task data.\"}]",1784211697,76,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"rl-forgets-towards-continual-policy-optimization","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/rl-forgets-towards-continual-policy-optimization/86431/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the paper address in continual post-training for vision-language models?","Question",{"text":76,"@type":77},"The paper addresses catastrophic forgetting, where learning new tasks during continual post-training degrades previously acquired capabilities.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"Why do the authors argue that existing evidence for “RL forgets less” is incomplete?",{"text":81,"@type":77},"Existing evaluations can underestimate forgetting due to outdated/homogeneous benchmarks and narrow task distributions that do not reflect modern multimodal reasoning workloads.",{"name":83,"@type":74,"acceptedAnswer":84},"How does Continual Policy Optimization (CPO) reduce forgetting differently from standard policy optimization?",{"text":85,"@type":77},"CPO replaces an intractable historical constraint with a replay-free framework based on a prior-task behavioral KL objective, limiting policy drift via sparse parameter-movement regularization rather than evaluating KL on current-task data.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":22,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":20,"slug":137},19,"General","general"]