[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86400-en":3,"doc-seo-86400-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86400,1649267921044,"Ava Thompson","https://us-avatar.wpscdn.com/avatar/1800007509477c92dfb?_k=1782875107921204101",8,"Research & Report","Attribution Guided Continual Learning for Large Language Models","Large language models face catastrophic forgetting in continual learning: after sequentially learning new tasks, performance on earlier tasks degrades. Prior techniques such as data replay, parameter freezing, and regularization reduce forgetting but do not explain which internal parameters store prior knowledge versus which can be updated for new tasks. This work introduces an attribution-guided continual fine-tuning framework using Layer-wise Relevance Propagation (LRP) to estimate parameter importance and constrain updates to critical parameters, reducing forgetting while maintaining adaptability.","arXiv :2605 .05285v2 [ cs .LG] 13 Jul 2026  \nAttribution-Guided Continual Learning for Large  \nLanguage Models  \nYazheng Liu 1 , Yuxuan Wan 1 , Rui Xu 1 , Xi Zhang2 , Sihong Xie 1 , Hui Xiong 1  \n1 The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China  \n2 The Beijing University of Posts and Telecommunications, Beijing, China  \nAbstract  \nLarge language models (LLMs) often suffer from catastrophic forgetting in continual learning: after learning new tasks sequentially, they perform worse on earlier tasks. Existing methods mitigate catastrophic forgetting by data replay, parameter freezing, or regularization. However, these methods lack understanding of LLM mechanisms and cannot distinguish which parameters store important knowledge from previous tasks and which parameters can be updated for new tasks. To address this, we propose the attribution-guided continual fine-tuning framework that leverages Layer-wise Relevance Propagation (LRP) to estimate parameter importance based on the internal computational process of LLMs. During continual learning, parameters critical to previous tasks are constrained to receive smaller updates, while less relevant parameters remain available for learning new tasks. Extensive experiments show that, compared with baseline methods, our approach reduces catastrophic forgetting while preserving adaptability to new tasks, highlighting the value of mechanistic attribution for continual fine-tuning of LLMs.  \n1 Introduction  \nLarge Language Models (LLMs) [1, 2, 3, 4, 5] have achieved exceptional performance across diverse tasks, such as multi-step reasoning[6, 7], instruction following [8, 9], and code geneartion [10, 11] . However, LLMs are usually trained on static, broad-domain data. This can lead to performance degradation when the target domain changes [12, 13, 14] . To adapt LLMs to downstream tasks while preserving prior knowledge, researchers use continual learning method [15, 16] . Continual learning trains models on a sequence of tasks but faces a challenge known as catastrophic forgetting [17, 18]: fine-tuning LLM on a new task can substantially degrade its performance on previously learned tasks.  \nExisting approaches have been proposed to mitigate catastrophic forgetting in continual learning. Replay methods [19, 20, 21, 22] retain examples from previous tasks and train them together with current task data. Regularization methods [23, 24, 25] penalize large deviations from the previous model in parameter space. Freezing methods fix selected LLM parameters to reduce forgetting [26] . However, these methods lack a mechanistic understanding of how knowledge are distributed across LLM parameters. They may constrain parameters that are necessary for adapting to new tasks or modify parameters that are critical for retaining prior knowledge.  \nTo understand the internal mechanisms of LLMs in continual learning, we first explore the importance of parameters in LLMs on different tasks. Let θ denote the model parameters, W (l)(θ) be the set of model parameters in its l-th layer, and W (l) ∈ W (l)(θ) represent an individual model parameter. For a task T, we compute the importance of each parameter with respect to the next-token logits using the attribution method introduced in Section 4.1.2 . We then select the top-K most important elements in W (l), denoted by PK 􀀀 W (l);θ, T 􀀁 . Given two tasks, T1 and T2 , we compare independent single-task fine-tuning with sequential continual fine-tuning, as illustrated in Figure 3(a,b) . In the single-task setting, a pretrained LLM is fine-tuned separately on each task, producing task-specific  \nPreprint.  \nmodels θ′1 and θ′2. We compute PK 􀀀 W (l); θ′1 , T1 􀀁and PK 􀀀 W (l); θ′2 , T2 􀀁 . In the continual setting, thepretrained model is fine-tuned sequentially on T1 and T2 , resulting in final parameters θ2. We obtain PK 􀀀 W (l);θ2 , T1 􀀁and PK 􀀀 W (l);θ2 , T2 􀀁 . Finally, we quantify the similarity of PK 􀀀 W (l); θ′1 , T1 􀀁 and PK 􀀀 W ","cbCaitgxPap473WK","https://ap.wps.com/l/cbCaitgxPap473WK","pdf",2590528,6,1,12,"English","en",105,"# Introduction\n## Problem: Catastrophic forgetting in continual learning\n## Existing methods and limitations\n## Parameter importance via attribution and LRP","[{\"question\":\"What problem does the paper address in continual learning for LLMs?\",\"answer\":\"It addresses catastrophic forgetting, where fine-tuning on a new task sequentially causes worse performance on previously learned tasks.\"},{\"question\":\"How does the proposed method identify parameters important to earlier tasks?\",\"answer\":\"It uses Layer-wise Relevance Propagation (LRP) to compute parameter importance based on the internal computations that affect next-token logits.\"},{\"question\":\"How does attribution guidance help prevent forgetting while enabling learning new tasks?\",\"answer\":\"The method constrains smaller updates for parameters deemed critical for old tasks, while allowing less relevant parameters to change for learning new tasks.\"}]",1784211519,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"attribution-guided-continual-learning-for-large-language-models","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/attribution-guided-continual-learning-for-large-language-models/86400/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the paper address in continual learning for LLMs?","Question",{"text":76,"@type":77},"It addresses catastrophic forgetting, where fine-tuning on a new task sequentially causes worse performance on previously learned tasks.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the proposed method identify parameters important to earlier tasks?",{"text":81,"@type":77},"It uses Layer-wise Relevance Propagation (LRP) to compute parameter importance based on the internal computations that affect next-token logits.",{"name":83,"@type":74,"acceptedAnswer":84},"How does attribution guidance help prevent forgetting while enabling learning new tasks?",{"text":85,"@type":77},"The method constrains smaller updates for parameters deemed critical for old tasks, while allowing less relevant parameters to change for learning new tasks.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,115,120,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":107,"slug":137},19,"General","general"]