[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122100-en":3,"doc-seo-122100-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122100,8796095461610,"Oliver","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Continual Task Allocation in Meta-Policy Network via Sparse Prompting","How to train a generalizable meta-policy by continually learning a sequence of tasks, while achieving both plasticity (fast adaptation) and stability (retaining common knowledge), remains difficult in reinforcement learning. CoTASP addresses this with continual task allocation via sparse prompting: it learns over-complete dictionaries to generate sparse binary masks that extract task-specific sub-networks from a meta-policy network. It optimizes prompts and sub-network weights alternately, then updates the dictionaries to align prompts with task embeddings. Experiments show strong plasticity-stability trade-offs without storing or replaying past experience.","Continual Task Allocation in Meta-Policy Network via Sparse Prompting  \nYijun Yang 1 2 Tianyi Zhou 3 Jing Jiang 2 Guodong Long 2 Yuhui Shi 1  \nAbstract  \nHow to train a generalizable meta-policy by continually learning a sequence of tasks? It is a natural human skill yet challenging to achieve by current reinforcement learning: the agent is expected to quickly adapt to new tasks (plasticity) meanwhile retaining the common knowledge from previous tasks (stability) . We address it by“Continual Task Allocation via Sparse Prompting (CoTASP)”, which learns over-complete dictionaries to produce sparse masks as prompts extracting a sub-network for each task from a meta-policy network. CoTASP trains a policy for each task by optimizing the prompts and the subnetwork weights alternatively. The dictionary is then updated to align the optimized prompts with tasks’ embedding, thereby capturing tasks’ semantic correlations. Hence, relevant tasks share more neurons in the meta-policy network due to similar prompts while cross-task interference causing forgetting is effectively restrained. Given a metapolicy and dictionaries trained on previous tasks, new task adaptation reduces to highly efficient sparse prompting and sub-network finetuning.  \nIn experiments, CoTASP achieves a promising plasticity-stability trade-off without storing or replaying any past tasks’ experiences. It outperforms existing continual and multi-task RL methods on all seen tasks, forgetting reduction, and generalization to unseen tasks. Our code is available at [https://github.com/stevenyangyj/CoTASP](https://github.com/stevenyangyj/CoTASP)  \n1. Introduction  \nAlthough reinforcement learning (RL) has demonstrated excellent performance on learning a single task, e.g., playing Go (Silver et al., 2016), robotic control (Schulman  \n1 Southern University of Science and Technology 2University of Technology Sydney 3University of Maryland, College Park. Correspondence to: Tianyi Zhou \u003C[tianyi@umd.edu](tianyi@umd.edu) >, Jing Jiang  \n\u003C[jing.jiang@uts.edu.au](jing.jiang@uts.edu.au) >, Yuhui Shi \u003C[shiyh@sustech.edu.cn](shiyh@sustech.edu.cn) >.  \nProceedings of the 40 th International Conference on Machine Learning, Honolulu, Hawaii, USA. PMLR 202, 2023 . Copyright 2023 by the author(s) .  \nSparse  \ncoding of   \nSparse prompts  \nFigure 1: Main steps and components of CoTASP.  \net al., 2017 ; Degrave et al., 2022), and offline policy optimization (Yu et al., 2020 ; Yang et al., 2022), it still suffers from catastrophic forgetting and cross-task interference when learning a stream of tasks on the fly (McCloskey & Cohen, 1989 ; Bengio et al., 2020) or a curated curriculum of tasks (Fang et al., 2019 ; Ao et al., 2021 ; 2022) . So it is challenging to train a meta-policy that can generalize to all learned tasks or even unseen ones with fast adaptation, which however is an inherent skill of human learning. This problem has been recognized as continual or lifelong RL (Mendez & Eaton, 2022) and attracted growing interest in recent RL research.  \nA primary and long-standing challenge in continual RL is the plasticity-stability trade-off (Khetarpal et al., 2022): the RL policy on the one hand needs to retain and reuse the knowledge shared across different tasks in history (stability) while on the other hand can be quickly adapted to new tasks without interference from previous tasks (plasticity) . Addressing this challenge is vital to improving the efficiency of  \ncontinual RL and the generalization capability of its learned policy. A meta-policy with better stability can reduce the necessity of experience replay and its memory/computation cost. Moreover, the required network size can be effectively reduced if the meta-policy can manage knowledge sharing across tasks in a more compact and efficient manner. Hence, stability can greatly improve the efficiency of continual RL when the number of tasks increases (Shin et al., 2017 ; Li & Hoiem, 2018) . Furthermore, better plasticity indicates f","cbCaiqm2uCA3AQYL","https://ap.wps.com/l/cbCaiqm2uCA3AQYL","pdf",4700359,1,16,"English","en",105,"# Introduction\n## Plasticity-stability trade-off in continual RL\n## Prompting-based meta-policy extraction\n## CoTASP: sparse prompts and dictionary learning","[{\"question\":\"What is the core idea of CoTASP for continual task learning?\",\"answer\":\"CoTASP learns layer-wise dictionaries to produce sparse binary masks (prompts) from task embeddings, then uses those masks to extract task-specific sub-networks from a meta-policy network.\"},{\"question\":\"How does CoTASP balance plasticity and stability?\",\"answer\":\"Relevant tasks reuse more neurons because their prompts are similar, enabling fast adaptation (plasticity), while irrelevant tasks share fewer or no neurons, restraining interference and forgetting (stability).\"},{\"question\":\"Does CoTASP rely on storing or replaying past task experiences?\",\"answer\":\"No. The method achieves a strong plasticity-stability trade-off without storing or replaying any previous tasks’ experiences.\"}]","Continual Task Allocation in Meta-Policy Network via Sparse Prompting | PDF",1785808825,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"continual-task-allocation-in-meta-policy-network-via-sparse-prompting","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/continual-task-allocation-in-meta-policy-network-via-sparse-prompting/122100/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the core idea of CoTASP for continual task learning?","Question",{"text":75,"@type":76},"CoTASP learns layer-wise dictionaries to produce sparse binary masks (prompts) from task embeddings, then uses those masks to extract task-specific sub-networks from a meta-policy network.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does CoTASP balance plasticity and stability?",{"text":80,"@type":76},"Relevant tasks reuse more neurons because their prompts are similar, enabling fast adaptation (plasticity), while irrelevant tasks share fewer or no neurons, restraining interference and forgetting (stability).",{"name":82,"@type":73,"acceptedAnswer":83},"Does CoTASP rely on storing or replaying past task experiences?",{"text":84,"@type":76},"No. The method achieves a strong plasticity-stability trade-off without storing or replaying any previous tasks’ experiences.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":29,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]