[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122168-en":3,"doc-seo-122168-105":31,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},122168,7971461740886,"Theodore","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","Does Continual Learning Equally Forget All Parameters? - Forgetting Prioritized Finetuning（FPF）与k-FPF方法","Distribution shift in continual learning (CL) typically causes catastrophic forgetting as models overwrite previously learned weights. While experience replay can alleviate this via buffered data, every-step replay is costly. This study analyzes training dynamics to identify which neural modules are most prone to forgetting, finding forgetting concentrated in a small set of task-sensitive modules. Finetuning only those modules with a small final buffer yields efficient gains via Forgetting Prioritized Finetuning (FPF). A lazy variant, k-FPF, removes frequent replay and triggers FPF only k times, achieving comparable accuracy with much lower overhead and cost across class- and domain-incremental benchmarks.","Does Continual Learning Equally Forget All Parameters?  \nHaiyan Zhao 1 Tianyi Zhou 2 Guodong Long 1 Jing Jiang 1 Chengqi Zhang 1  \nAbstract  \nDistribution shift (e.g., task or domain shift) in continual learning (CL) usually results in catastrophic forgetting of previously learned knowledge. Although it can be alleviated by repeatedly replaying buffered data, the every-step replay is time-consuming. In this paper, we study which modules in neural networks are more prone to forgetting by investigating their training dynamics during CL. Our proposed metrics show that only a few modules are more task-specific and sensitive to task change, while others can be shared across tasks as common knowledge. Hence, we attribute forgetting mainly to the former and find that finetuning them only on a small buffer at the end of any CL method can bring non-trivial improvement. Due to the small number of finetuned parameters, such “Forgetting Prioritized Finetuning (FPF)” is efficient in computation. We further propose a more efficient and simpler method that entirely removes the every-step replay and replaces them by only k-times of FPF periodically triggered during CL. Surprisingly, this “kFPF” performs comparably to FPF and outperforms the SOTA CL methods but significantly reduces their computational overhead and cost. In experiments on several benchmarks of classand domain-incremental CL, FPF consistently improves existing CL methods by a large margin, and k-FPF further excels in efficiency without degrading the accuracy. We also empirically studied the impact of buffer size, epochs per task, and finetuning modules on the cost and accuracy of our methods.  \n1University of Technology Sydney 2University of Maryland. Correspondence to: Haiyan Zhao \u003CHaiyan.Zhao- [2@student.uts.edu.au](2@student.uts.edu.au) >, Tianyi Zhou \u003C[tianyi@umd.edu](tianyi@umd.edu) >, Guodong Long, Jing Jiang, Chengqi Zhang \u003C{guodong.long, jing.jiang, [chengqi.zhang](chengqi.zhang}@uts.edu.au)[}](chengqi.zhang}@uts.edu.au)[@uts.edu.au](chengqi.zhang}@uts.edu.au)>.  \nProceedings of the 40 th International Conference on Machine Learning, Honolulu, Hawaii, USA. PMLR 202, 2023 . Copyright 2023 by the author(s) .  \n1. Introduction  \nEmpowered by advancing deep learning techniques and neural networks, machine learning has achieved unprecedented promising performance on challenging tasks in different fields, mostly under the i.i.d. offline setting. However, its reliability and performance degenerate drastically in continual learning (CL) where the data distribution or task in training changes over time, as the model quickly adapts toa new task and overwrites the previously learned weights. This leads to a severe bias toward more recent tasks and“catastrophic forgetting” of previously learned knowledge, which is detrimental to a variety of practical applications.  \nA widely studied strategy to mitigate forgetting is experience replay (ER) (Ratcliff, 1990 ; Robins, 1995) and its variants (Riemer et al., 2018 ; Buzzega et al., 2020 ; Boschini et al., 2022), which store a few data from previous tasks in the limited memory and train the model using both the current and buffered data. However, they only bring marginal improvements when the memory is too small to store sufficient data for recovering previously learned knowledge, which is common due to the complicated distributions of previous tasks. In contrast, multi-task learning (Caruana, 1997) usually adopts a model architecture composed of atask-agnostic backbone network and multiple task-specific adapters on top of it. While the backbone needs to be pre-trained on large-scale data, the adapters are usually lightweight and can be achieved using a few data. In CL, however, we cannot explicitly pre-define and separate the task-agnostic parts and task-specific parts. Although previous methods (Schwarz et al., 2018 ; Zenke et al., 2017) have studied to restrict the change of parameters critical to previous tasks, such an extra constra","cbCaiiiodH9vYf04","https://ap.wps.com/l/cbCaiiiodH9vYf04","pdf",1116595,2,1,24,"English","en",105,"# Abstract\n# Introduction\n## Continual learning and catastrophic forgetting\n## Experience replay and its limitations\n## Task-specific adapters and plasticity-stability trade-off\n## Research questions and proposed approach","[{\"question\":\"Why does continual learning suffer catastrophic forgetting under distribution shift?\",\"answer\":\"As tasks or data distributions change over time, models quickly adapt to new tasks and overwrite weights learned for earlier tasks, creating a strong bias toward recent tasks.\"},{\"question\":\"How does this work identify which parameters cause forgetting?\",\"answer\":\"It investigates training dynamics by measuring how model parameters change over time during continual learning, revealing that only a few modules change more drastically between tasks.\"},{\"question\":\"What are FPF and k-FPF, and how do they reduce computational cost?\",\"answer\":\"FPF finetunes task-specific parameters using buffered data only at the end of continual learning methods. k-FPF further removes every-step replay by triggering FPF periodically only k times, reducing overhead while maintaining comparable accuracy.\"}]","Does Continual Learning Equally Forget All Parameters? - Forgetting Prioritized Finetuning（FPF）与k-FPF方法 | PDF",1785809161,60,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":29},"does-continual-learning-equally-forget-all-parameters-forgetting-prioritized-finetuning-fpf-and-k-fpf","",{"@graph":37,"@context":85},[38,54,68],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,48,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":20},"https://docshare.wps.com/document/","Document",{"item":49,"name":12,"@type":44,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/does-continual-learning-equally-forget-all-parameters-forgetting-prioritized-finetuning-fpf-and-k-fpf/122168/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why does continual learning suffer catastrophic forgetting under distribution shift?","Question",{"text":75,"@type":76},"As tasks or data distributions change over time, models quickly adapt to new tasks and overwrite weights learned for earlier tasks, creating a strong bias toward recent tasks.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does this work identify which parameters cause forgetting?",{"text":80,"@type":76},"It investigates training dynamics by measuring how model parameters change over time during continual learning, revealing that only a few modules change more drastically between tasks.",{"name":82,"@type":73,"acceptedAnswer":83},"What are FPF and k-FPF, and how do they reduce computational cost?",{"text":84,"@type":76},"FPF finetunes task-specific parameters using buffered data only at the end of continual learning methods. k-FPF further removes every-step replay by triggering FPF periodically only k times, reducing overhead while maintaining comparable accuracy.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":47,"category_name":107,"show_sort_weight":30,"slug":108},5,"Comic","comic",{"id":110,"doc_module":4,"doc_module_name":47,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":47,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":47,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":47,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":47,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]