[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81818-en":3,"doc-seo-81818-105":31,"detail-sidebar-cat-0-en-105":93},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},81818,4398048950312,"Violet","https://ap-avatar.wpscdn.com/avatar/400002538284de19e3c?_k=1778320343897328908",8,"Research & Report","Procedural Memory Distillation Online Reflection for Self-Improving Language Models","Reinforcement learning with verifiable rewards (RLVR) and self-distillation variants evaluate each rollout via a verifier and update the policy using an episode-level signal. Yet they rarely preserve the richer procedural information within rollouts, even though repeated encounters with related problems across epochs create cross-episode signals about consistently passing strategies, persistent failure modes, and recurring patterns. Procedural Memory Distillation (PMD) converts these signals into reusable procedural memory extracted online and distills it into policy weights during training, producing a memory-free model at inference.","arXiv :2607 .0 1480v 1 [ cs .AI] 1 Jul 2026  \nProcedural Memory Distillation: Online Reflection for Self-Improving Language Models  \nYe Liu, Srijan Bansal, Bo Pang, Yang Li,  \nZeyu Leo Liu, Yifei Ming, Zixuan Ke, Shafiq Joty, Semih Yavuz  \nSalesforce AI Research  \nyeliu, srijanbansal, sjoty, [syavuz @salesforce.com](syavuz @salesforce.com)  \nAbstract  \nReinforcement learning with verifiable rewards (RLVR), along with recent selfdistillation variants such as SDPO, evaluates each rollout against a verifier and updates the policy from that episode-level signal. However, the richer procedural information in the rollout is rarely retained or reused. Across episodes and epochs, the model repeatedly encounters related problems under a changing policy, producing cross-episode signals that episode-local updates cannot capture: which strategies consistently pass verification, which failure modes persist, which patterns recur.  \nWe propose Procedural Memory Distillation (PMD), which converts these crossepisode signals into reusable procedural memory and distills it into the policy’s weights during training. This memory functions as a training scaffold, absorbed into the policy itself, yielding a memory-free model at inference. PMD organizes the memory at three levels of abstraction: raw trajectories, self-reflected strategies and lessons, and higher-level behavioral patterns that recur across problems, all extracted online from the model’s own trajectories. A memory-conditioned self-teacher draws on the accumulated experience to supervise the student on its own rollouts, enabling student to progressively internalize procedural knowledge within its parameters. The central design principle is co-evolution: the policy generates rollouts that update the memory, and memory shapes the supervision that updates the policy. Empirically, across Qwen3-8B and OLMo3-Instruct-7B, PMD improves over SDPO by 3.8–5.5% on SCIKNOWEVAL and 7.9–13.6% on LIVECODEBENCH. Co-evolution powers these gains: freezing either the memory or the policy trails PMD by more than 10% across SCIKNOWEVAL domains.  \n1 Introduction  \nThe prevailing paradigm for preference optimization and reinforcement learning with verifiable rewards (RLVR) operates at the level of individual episodes; methods such as PPO [25], DPO [22] and GRPO [26, 6] convert per-rollout preferences or outcome-based checks into learning signals. Each rollout receives a reward, feedback or hindsight correction; the policy gets updated accordingly; and the experience is discarded. This design is natural when episodes are independent. However, in practice, models repeatedly encounter the same or related problems across epochs under a continually evolving policy. These repeated interactions carry cross-episode signals that isolated, one-step updates cannot capture: which strategies pass verification, which failure modes persist, which patterns recur.  \nRecent work suggests that the learning signal available during training can be richer than a scalar reward alone. For instance, in Self-Distillation Policy Optimization (SDPO) [10], the current policy can act as self-teacher when conditioned on training-time context: textual feedback when available, or a successful sibling rollout from the same group when it is not. This reflects a broader shift from offline distillation to on-policy distillation, where the student is trained on states it actually visits rather than static teacher demonstrations [8, 23, 1, 5, 30] .  \nPreprint.  \nWhile SDPO and related on-policy distillation approaches [1, 53] help alleviate the sparse reward limitations of standard RLVR and the distributional mismatch of offline distillation, their updates remain episode-local: they do not systematically preserve what the model has discovered across earlier attempts. We propose Procedural Memory Distillation (PMD), which converts these repeated attempts into reusable procedural memory and distills it into the policy’s weights during training. PMD b","cbCaisn3zmtCtdx2","https://ap.wps.com/l/cbCaisn3zmtCtdx2","pdf",1124172,5,1,24,"English","en",105,"# Introduction\n## Episode-local learning and cross-episode signals\n## Procedural Memory Distillation (PMD)\n### Co-evolution of policy and memory\n### Three-level procedural memory hierarchy\n## Empirical results","[{\"question\":\"What limitation do RLVR and SDPO face in learning from rollout information?\",\"answer\":\"They update policies using episode-local verifier signals, while the richer procedural information contained in rollouts is not systematically preserved or reused across episodes and epochs.\"},{\"question\":\"How does PMD create reusable procedural memory during training?\",\"answer\":\"PMD converts cross-episode procedural signals into an online procedural memory and distills it into the policy’s weights, using a memory-conditioned self-teacher to supervise the student on its own rollouts.\"},{\"question\":\"Why is co-evolution between the policy and memory important in PMD?\",\"answer\":\"The policy generates rollouts that update the memory, and the updated memory shapes supervision that trains the next policy version; freezing either component breaks this alignment and reduces performance by more than 10%.\"}]","Procedural Memory Distillation Online Reflection for Self-Improving Language Models | PDF",1784176342,60,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":88,"head_meta":90,"extra_data":92,"updated_unix":29},"procedural-memory-distillation-online-reflection-for-self-improving-language-models","",{"@graph":37,"@context":87},[38,55,70],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":54},"https://docshare.wps.com/document/procedural-memory-distillation-online-reflection-for-self-improving-language-models/81818/",4,{"url":53,"name":13,"@type":56,"author":57,"headline":13,"publisher":59,"fileFormat":62,"inLanguage":24,"description":14,"dateModified":63,"datePublished":64,"encodingFormat":62,"isAccessibleForFree":65,"interactionStatistic":66},"DigitalDocument",{"name":9,"@type":58},"Person",{"url":42,"name":60,"@type":61},"DocShare","Organization","application/pdf","2026-07-29","2026-07-16",true,{"@type":67,"interactionType":68,"userInteractionCount":20},"InteractionCounter",{"@type":69},"ViewAction",{"@type":71,"mainEntity":72},"FAQPage",[73,79,83],{"name":74,"@type":75,"acceptedAnswer":76},"What limitation do RLVR and SDPO face in learning from rollout information?","Question",{"text":77,"@type":78},"They update policies using episode-local verifier signals, while the richer procedural information contained in rollouts is not systematically preserved or reused across episodes and epochs.","Answer",{"name":80,"@type":75,"acceptedAnswer":81},"How does PMD create reusable procedural memory during training?",{"text":82,"@type":78},"PMD converts cross-episode procedural signals into an online procedural memory and distills it into the policy’s weights, using a memory-conditioned self-teacher to supervise the student on its own rollouts.",{"name":84,"@type":75,"acceptedAnswer":85},"Why is co-evolution between the policy and memory important in PMD?",{"text":86,"@type":78},"The policy generates rollouts that update the memory, and the updated memory shapes supervision that trains the next policy version; freezing either component breaks this alignment and reduces performance by more than 10%.","https://schema.org",{"og:url":53,"og:type":89,"og:title":13,"og:site_name":60,"og:description":14},"article",{"robots":91,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":94},[95,99,103,107,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":96,"show_sort_weight":97,"slug":98},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},"Literature",80,"literature",{"id":54,"doc_module":4,"doc_module_name":47,"category_name":104,"show_sort_weight":105,"slug":106},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":30,"slug":109},"Comic","comic",{"id":111,"doc_module":4,"doc_module_name":47,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":47,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":47,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":47,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":47,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":47,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]