[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81665-en":3,"doc-seo-81665-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81665,16904993612988,"Olivia Brown","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories","Language Models Need Sleep proposes a continual-learning paradigm for large language models that addresses their inherent static knowledge after deployment. The approach introduces a “Sleep” stage that consolidates short-term, fragile in-context memories into stable long-term parameters. Sleep combines Knowledge Seeding via upward distillation into a larger network and a “Dreaming” phase that uses RL to create synthetic-data curricula for rehearsal and improvement without human supervision. Experiments on long-horizon continual learning support the value of Sleep for knowledge incorporation and few-shot generalization.","arXiv :2606 .03979v2 [ cs .LG] 10 Jul 2026  \nLanguage Models Need Sleep: Learning to Self-Modify and Consolidate Memories  \nAli Behrouz†, Farnoosh Hashemi‡, Adel Javanmard †, and Vahab Mirrokni †  \n†  ‡  \nAbstract  \nThe past few decades have witnessed significant advances in the design of machine learning algorithms–from early studies on task-specific shallow models to more general deep Large Language Models (LLMs) . Despite showing promising results in tasks that require instant prediction or in-context learning, existing models lack the ability to continually learn and effectively transfer their temporal in-context knowledge to their long-term parameters. Inspired by human learning process, we introduce a “Sleep” paradigm that allows the models to continually learn, distill their short-term fragile memories into stable long-term knowledge with replay, and recursively improve themselves with “Dreaming” process. In more detail, sleep consists of two stages: (1) Memory Consolidation: an upward distillation process, called Knowledge Seeding, where the memories of a smaller-self are distilled into a larger network to provide more capacity while preserving the knowledge. As a proof of concept, we present a new Generalized Distillation process for Knowledge Seeding (i.e., the combination of on-policy distillation with Reinforcement Learning (RL)-based imitation learning); (2) Dreaming: a self-improvement phase, where the model uses RL to generate a curriculum of synthetic data to rehearse new knowledge and refine existing capabilities without human supervision. Our experiments on long-horizon, continual learning, knowledge incorporation, and few-shot generalization tasks support the importance of the sleep stage.  \n1 Introduction  \nThe development of Large Language Models (LLMs) marks a pivotal milestone in machine learning research: a paradigm shift from task-specific models to more general-purpose systems with various emergent capabilities (Brown et al. 2020; Schaeffer et al. 2023) . Despite LLMs’ remarkable capabilities in diverse sets of tasks (Nijkamp et al. 2023; Wang et al. 2023; Comanici et al. 2025), they are largely static after their initial deployment, meaning that they successfully perform tasks learned during pre-or post-training, but are unable to continually acquire new capabilities beyond their immediate context. This inherent static nature creates a crucial vulnerability: The model’s knowledge and skills become progressively stale, operating with a fixed \"knowledge cutoff\" date beyond which it is unaware of new facts, events, and evolving information (Cheng et al. 2024) .  \nEfforts to overcome this limitation have primarily focused on: (1) re-pretraining on an expanded dataset, which despite its effectiveness, is computationally expensive and impractical for frequent updates (Ibrahim et al. 2024); (2) using expensive continual parameter updates or other lightweight alternatives, such as fine-tuning or low-rank adaption (Hu et al. 2022; Akyürek et al. 2024a), which with iterative updates often results in Catastrophic Forgetting (CF) (Kemker et al. 2018; Shi et al. 2024)–a well-known phenomenon where the model’s proficiency on original tasks degrades catastrophically as it learns new ones. This dilemma—between knowledge obsolescence on one hand and catastrophic forgetting as well as the prohibitive cost or destructive nature of updates on the other—underscores a critical, unresolved challenge: enabling LLMs to learn incrementally and efficiently throughout their lifecycle.  \nIn recent years, In-Context Learning (ICL) (Brown et al. 2020) has gained attention as a highly efficient and successful form of continual learning (Akyürek et al. 2022, 2024b; Dong et al. 2024; Li et al. 2025) . Initially, ICL was known as an emergent ability of LLMs that is trained on large scale data, enabling them to adapt fast to the context and so perform zero-or few-shot tasks (Brown et al. 2020) . Later, more studies revealed and formali","cbCaiqmKLCZDDc8x","https://ap.wps.com/l/cbCaiqmKLCZDDc8x","pdf",3151827,2,1,26,"English","en",105,"# Introduction\n## Memory Staleness in Static LLMs\n## Continual Learning Challenges\n## In-Context Learning as Continual Learning\n## Sleep Paradigm: Knowledge Seeding and Dreaming","[{\"question\":\"Why do existing language models struggle with continual learning after deployment?\",\"answer\":\"They are largely static after initial deployment, so knowledge becomes stale beyond a fixed knowledge cutoff and the model cannot continually acquire new capabilities outside the immediate context window.\"},{\"question\":\"What is the “Sleep” paradigm proposed in the document?\",\"answer\":\"Sleep is a two-stage process that helps models continually learn by consolidating short-term fragile memories into stable long-term knowledge, combining Memory Consolidation (Knowledge Seeding) and a self-improvement “Dreaming” phase.\"},{\"question\":\"How do Knowledge Seeding and Dreaming work together?\",\"answer\":\"Knowledge Seeding performs upward distillation from a smaller model into a larger network to preserve knowledge while adding capacity. Dreaming then uses reinforcement learning to generate synthetic-data curricula for rehearsal, refining existing capabilities without human supervision.\"}]",1784175293,66,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"language-models-need-sleep-learning-to-self-modify-and-consolidate-memories","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/language-models-need-sleep-learning-to-self-modify-and-consolidate-memories/81665/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do existing language models struggle with continual learning after deployment?","Question",{"text":75,"@type":76},"They are largely static after initial deployment, so knowledge becomes stale beyond a fixed knowledge cutoff and the model cannot continually acquire new capabilities outside the immediate context window.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the “Sleep” paradigm proposed in the document?",{"text":80,"@type":76},"Sleep is a two-stage process that helps models continually learn by consolidating short-term fragile memories into stable long-term knowledge, combining Memory Consolidation (Knowledge Seeding) and a self-improvement “Dreaming” phase.",{"name":82,"@type":73,"acceptedAnswer":83},"How do Knowledge Seeding and Dreaming work together?",{"text":84,"@type":76},"Knowledge Seeding performs upward distillation from a smaller model into a larger network to preserve knowledge while adding capacity. Dreaming then uses reinforcement learning to generate synthetic-data curricula for rehearsal, refining existing capabilities without human supervision.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]