[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86109-en":3,"doc-seo-86109-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86109,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Can a Language Model Learn Facts Continually in Its Weights","Continual learning aims to let a language model acquire new factual knowledge after training by writing each fact into its weights. Experiments with invented facts inserted into Qwen3 models track performance from creation through 20–100 subsequent writes using held-out questions of five types. Training-data breadth shapes the stored knowledge: bare-statement training yields recitation, while diverse restatements greatly narrow the recitation-to-use gap. Later writes also cause behavioural forgetting without deletion; recovered accuracy depends on supplying the fact in context, while earlier facts remain largely unreachable.","Preprint, July 2026  \narXiv :2607 . 1 1020v 1 [ cs .CL] 13 Jul 2026  \nCan a Language Model Learn Facts Continually in Its Weights?  \nCharles O’Neill1  \n1 Baseten  \nAbstract  \nContinual learning promises a language model that keeps acquiring knowledge after training, with each new fact written into its weights. Whether weight writes can support accumulation remains undecided. We follow invented facts written into Qwen3 models from creation through sequences of twenty to one hundred later writes, using held-out questions of five types, with the original model given the fact in its prompt as the reference. Across these experiments, the breadth of the training data determines the kind of knowledge created. Bare-statement training produces recitation, while diverse restatements reduce the recitation-to-use gap from 27.4 to 5.4 points without showing the model a conclusion. This difference carries into later writes: after twenty sequential writes, bare-statement facts retain 1% accuracy while facts written from broad study data retain 46% . We also find that facts can be behaviourally forgotten without being erased. Forgotten facts keep most of the log-probability added by their write, and under bare-statement training 70% of wrong answers about them contain the most recently written fact. The same writes barely degrade the model’s use of facts in context, and a forgotten study fact supplied in the prompt recovers to 77–80% on its questions. These results describe knowledge that is stored but question-keyed: later writes redirect the questions that reached it. Damage to unrelated abilities tracks KL divergence from the original model, and the later writes cause interference regardless of how the earlier fact was stored. Broad data can create usable knowledge, and a frozen reference can preserve capability, but no intervention we tested, including those built on accurate local measurements of each write, keeps earlier facts reachable. When facts must be composed or survive later writes, the reliable channel is context rather than the weights.  \n1. Introduction  \nA language model can hold a new fact in its context or its weights. Context makes the fact immediately usable, but only for the life of the prompt. Continual learning asks the weights to acquire facts after training and retain them through further writes. To understand whether they can, we characterise the object that a write creates: what kind of knowledge it contains, whether later writes preserve it, and what remains after questions about it fail. This turns catastrophic forgetting into a property of the written object, measured against the same content placed in context.  \nFine-tuning learns unknown facts slowly, and learned facts fail reversals and multi-hop use that the same facts support in a prompt (Gekhman et al., 2024; Berglund et al., 2024; Lampinen et al., 2025) . Diverse paraphrases make facts more extractable (Allen-Zhu and Li, 2024), while distributional drift predicts forgetting under further training (Shenfeld et al., 2025) . These observations lack a common account of why some writes produce usable knowledge, why some survive, and what a forgotten fact leaves behind.  \nWe build that account with invented facts, held-out questions of five types, and two fixed references: the original model and the same model with the fact in its prompt. Following each fact from creation through twenty to one hundred later writes lets us connect what a write creates to what later survives. Training-data breadth determines whether the model learns recitation or stated conclusions, and this difference predicts retention. When a fact eventually fails every question, checkpoint reconstructions still find most of its write’s log-probability lift in the weights, and under bare-statement writes the questions that once reached it return the newest write’s content instead. The same access problem is present before any overwriting: two individually usable written facts largely cannot be","cbCaiuqYoi3pLTUY","https://ap.wps.com/l/cbCaiuqYoi3pLTUY","pdf",904754,5,1,32,"English","en",105,"# Abstract\n# Introduction\n## Prior work and motivation\n## Experimental setup and evaluation approach\n## Contributions and key findings","[{\"question\":\"Can weight writing support continual accumulation of facts in language models?\",\"answer\":\"The results suggest weight writing can store knowledge, but accumulation that stays reliably usable is limited. Later writes often redirect questions keyed to earlier facts, making earlier facts hard to retrieve.\"},{\"question\":\"How does the breadth of training data affect what the model learns from a write?\",\"answer\":\"Broad training data produces knowledge that is more usable and less focused on recitation. Diverse restatements reduce the recitation-to-use gap substantially compared with bare-statement training.\"},{\"question\":\"If facts are forgotten in the weights, are they fully erased?\",\"answer\":\"No. Forgotten facts can be behaviourally forgotten without being erased; the write still contributes probability mass, but the model’s wrong answers tend to reflect the most recently written fact.\"}]",1784208583,81,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"can-a-language-model-learn-facts-continually-in-its-weights","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/can-a-language-model-learn-facts-continually-in-its-weights/86109/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Can weight writing support continual accumulation of facts in language models?","Question",{"text":76,"@type":77},"The results suggest weight writing can store knowledge, but accumulation that stays reliably usable is limited. Later writes often redirect questions keyed to earlier facts, making earlier facts hard to retrieve.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the breadth of training data affect what the model learns from a write?",{"text":81,"@type":77},"Broad training data produces knowledge that is more usable and less focused on recitation. Diverse restatements reduce the recitation-to-use gap substantially compared with bare-statement training.",{"name":83,"@type":74,"acceptedAnswer":84},"If facts are forgotten in the weights, are they fully erased?",{"text":85,"@type":77},"No. Forgotten facts can be behaviourally forgotten without being erased; the write still contributes probability mass, but the model’s wrong answers tend to reflect the most recently written fact.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]