[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82252-en":3,"doc-seo-82252-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82252,962075114765,"Quinn","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Interference and Retention in Continual Learning","Continual learning methods that rely on replay, elastic regularization, or distillation treat forgetting indirectly, after interference has already altered model behavior. The work models forgetting directly as interference between tasks, showing that in the frozen-feature regime task-A forgetting equals the interference energy induced by the update on the old task. For deep networks, a path-averaged curvature approximation recovers the same quantity with few extra forward passes, enabling structural criteria for lossless retention, non-zero distortion floors under conflicting overlap, and task-aware orthogonalization for optimal merging.","arXiv :2607 .09202v 1 [ cs .LG] 10 Jul 2026  \nInterference and Retention in Continual Learning  \nJulius Störk 1 ,∗  \n1VARTA Microbattery GmbH, Ellwangen, Germany  \n∗ Corresponding author: [julius.stoerk@varta-ag.com](julius.stoerk@varta-ag.com)  \n10.07.2026  \nAbstract  \nContinual learning commonly relies on post-hoc mechanisms such as replay, elastic regularization, or distillation. This work argues that forgetting should instead be modeled directly as interference between tasks. In the frozen-feature regime, forgetting from learning a new task is exactly the interference energy induced on the old task. In deep networks, the same quantity is recovered through path-averaged curvature with minimal additional forward passes.  \nWhen task supports are disjoint, forgetting can be eliminated structurally and when task supports overlap in conflicting directions, a non-zero distortion floor is unavoidable. The same geometry optimally merges models through task-aware orthogonalization. From this analysis we derive Interference-Gated Functional Allocation (igfa), a replay-free, Fisher-free method that shares directions when tasks align and protects them when they conflict. Across benchmarks, igfa achieves lossless retention when tasks are structurally separable and moves unavoidable cost from irreversible forgetting into deferred but recoverable plasticity when they are not. It matches the strongest replay-free structural baselines on dissimilar-task streams and improves on unconditional projection when similarity makes transfer worth preserving.  \nKeywords: continual learning; catastrophic forgetting; model merging; neural tangent kernel; gradient projection; rate–distortion; representation geometry  \n1 Introduction  \nA model trained sequentially on task B after task A often degrades on task A, a phenomenon known as catastrophic forgetting. Networks trained by stochastic gradient descent (SGD) or its variants update from the current minibatch alone (or a smoothed average of a short window of minibatches); the update is therefore oblivious to past knowledge. This obliviousness is desirable when the training data are i.i.d. , but harmful once the training distribution shifts over time in the continual setting. Standard approaches such as replay [3, 4 , 5], elastic weight consolidation [1, 2], and distillation [6] mostly repair forgetting after interference has already occurred. What is still missing is a predictive account of when interference is avoidable, when it is inevitable, and what structure an update must satisfy to retain old behavior without sacrificing useful transfer.  \nThis work focuses on the linear-on-features regime in which a frozen feature extractor ϕ (x) is paired with a trainable linear head, covering the common frozen-backbone and parameter-efficient fine-tuning (PEFT) settings and the first-order neural tangent [11] approximation of a fully trained network. The forgetting of an earlier task A after an update ∆ caused by learning task B is exactly the interference energy 12 ∆⊤ ΣA ∆ , where ΣA = Ex∼DA 􀀂ϕ(x)ϕ (x)⊤􀀃 denotes the feature second moment of task A, measuring which feature directions are active for that task. The same geometry predicts per-task forgetting accurately and extends approximately to deeper networks with feature drift.  \nThe paper makes three contributions.  \n1. Interference functional and removability. We derive an exact interference functional for forgetting. In the frozen-feature regime, forgetting is exactly the old task’s interference energy under the new update, which yields a clean criterion for when retention is lossless and when a distortion floor is unavoidable. The geometry of ΣA penalizes only components in its active subspace, while updates inker ΣA leave task A’s loss unchanged. This yields a structural criterion for lossless retention: interference is removable for disjoint task supports (reducing to ordinary orthogonality in isotropic geometry), whereas overlapping supports imply a n","cbCaie7PAj8lJC0f","https://ap.wps.com/l/cbCaie7PAj8lJC0f","pdf",3190334,1,41,"English","en",105,"# Introduction\n## Interference functional and removability\n## Optimal merging as Σ-orthogonalization","[{\"question\":\"Why does catastrophic forgetting occur in continual learning?\",\"answer\":\"Sequential SGD updates on task B use only current minibatch information, making the update oblivious to past knowledge. When the training distribution shifts over time, this harms performance on earlier task A.\"},{\"question\":\"How does the paper define forgetting in the frozen-feature regime?\",\"answer\":\"For an update Δ caused by learning task B, forgetting of task A equals the interference energy 1/2 · Δᵀ ΣA Δ, where ΣA is the feature second moment for task A.\"},{\"question\":\"When can forgetting be eliminated, and when is a distortion floor unavoidable?\",\"answer\":\"For disjoint task supports, forgetting can be removed structurally. When task supports overlap in conflicting directions, a non-zero distortion floor is unavoidable.\"}]",1784179177,103,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"interference-and-retention-in-continual-learning","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/interference-and-retention-in-continual-learning/82252/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why does catastrophic forgetting occur in continual learning?","Question",{"text":75,"@type":76},"Sequential SGD updates on task B use only current minibatch information, making the update oblivious to past knowledge. When the training distribution shifts over time, this harms performance on earlier task A.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the paper define forgetting in the frozen-feature regime?",{"text":80,"@type":76},"For an update Δ caused by learning task B, forgetting of task A equals the interference energy 1/2 · Δᵀ ΣA Δ, where ΣA is the feature second moment for task A.",{"name":82,"@type":73,"acceptedAnswer":83},"When can forgetting be eliminated, and when is a distortion floor unavoidable?",{"text":84,"@type":76},"For disjoint task supports, forgetting can be removed structurally. When task supports overlap in conflicting directions, a non-zero distortion floor is unavoidable.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]