[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84266-en":3,"doc-seo-84266-105":30,"detail-sidebar-cat-0-en-105":96},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84266,1374391974564,"Clementine","https://ap-avatar.wpscdn.com/avatar/14000253aa45c000a9e?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779874745381141002",8,"Research & Report","Co-LMLM: Continuous-Query Limited Memory Language Models","Limited memory language models (LMLMs) externalize factual knowledge into a human-readable knowledge base (KB) during pre-training, then retrieve needed facts during generation instead of storing them solely in model weights. Building on this, CO-LMLM replaces relational KB and explicit decoded queries with continuous vector queries and an unstructured index of vector keys paired with textual values. A free-form factual-span annotation pipeline removes prior Wikipedia-only constraints. Experiments on Wikipedia and FineWeb-Edu show improved perplexity and factual precision, with SimpleQA-verified results comparable to leading models at 360M scale.","arXiv :2607 .07707v 1 [ cs .CL] 8 Jul 2026  \nCo-LMLM: Continuous-Query Limited Memory  \nLanguage Models  \nYair Feldman* Linxi Zhao Nathan Godey Dongyoung Go Yilun Hua  \nKilian Q. Weinberger Jennifer J. Sun Yoav Artzi  \nDepartment of Computer Science  \nCornell University  \n[yairf@cs.cornell.edu](yairf@cs.cornell.edu)  \n{lz586, ng554, dg793, yh2228, kilian, jennifer.sun, [yoavartzi}@cornell.edu](yoavartzi}@cornell.edu)  \nAbstract  \nLimited memory language models (LMLMs) externalize factual knowledge during pre-training to a knowledge base (KB), rather than memorizing it in their weights.  \nDuring generation, the model then fetches knowledge from the KB as needed.  \nThis recently introduced paradigm provides multiple advantages, including knowledge control capabilities that remain beyond conventional LLMs. We propose continuous-query LMLM (CO-LMLM), where the KB pairs continuous keys with textual knowledge values, a significant departure from prior reliance on relational KB and queries. CO-LMLM generates flexible vector queries at minimal cost, while still integrating human-readable and attributable retrieved knowledge into its generation. We pair this design with an annotation pipeline that tags free-form factual spans in arbitrary text, removing prior work’s restriction to Wikipedia.  \nAcross pretraining on Wikipedia and FineWeb-Edu and at multiple model scales, CO-LMLM outperforms prior LMLMs and vanilla LLMs in both perplexity and factual precision. At 360M scale, this includes lower perplexity than models pretrained on 40× more data, and SimpleQA-verified performance that is in line with gpt-4o-mini and higher than Claude Sonnet 4.5 .  \n1 Introduction  \nRecently, there has been increasing interest in large language models (LLMs) that are trained to externalize knowledge [Ghosal et al., 2025, Zhao et al., 2026, Pouransari et al., 2026] . A particularly compelling approach is the limited memory language model [LMLM; Zhao et al., 2026], where the LLM is pre-trained to externalize knowledge into a human-readable knowledge base (KB) . This design offers several advantages. The KB is interpretable and easily editable, allowing for unlearning with no utility tradeoff, and enabling easy attribution of knowledge used to the source material.  \nZhao et al. instantiate the LMLM approach with a relational KB, which we refer to as REL-LMLM. Although representing a significant departure from conventional LLMs, REL-LMLM proposes a training process that follows the common scalable next-token-prediction pre-training, with the key difference of pre-processing the data to simulate knowledge retrieval. Experiments demonstrate significantly lower perplexities compared to conventional LLMs and factuality scores similar to models several orders of magnitude larger, while retaining similar utility (i.e., NLU scores) to LLMs of the same size.  \n*Individual author contributions are detailed in the acknowledgments.  \nPreprint.  \nFigure 1: Knowledge separation across three regimes. A standard LLM with RAG retrieves over external documents but keeps factual knowledge in its parameters (left); LMLM externalizes facts toa relational KB queried with an explicit decoded query (middle); CO-LMLM externalizes facts to an unstructured index, retrieved directly from the model’s hidden state (right) .  \nHowever, REL-LMLM has several key limitations that constrain the scaling of the pre-training process and the knowledge retrieval expressivity. Pre-training relies on Wikipedia, where each article is centered on a specific entity, making it straightforward to automatically annotate relational queries at scale during data pre-processing. Although Wikipedia is sufficient for a proof-of-concept demonstration, it provides no avenue to scale much further. The relational representation itself, although human readable, introduces several limitations. It restricts the data that can be externalized to items that are the object of natural language relational tuples, where bot","cbCaipkf24YLNQ1h","https://ap.wps.com/l/cbCaipkf24YLNQ1h","pdf",2017387,5,1,32,"English","en",105,"# Introduction\n## Limited memory language model (LMLM) paradigm\n## Relational LMLM limitations\n## Continuous-query LMLM (CO-LMLM) design\n## Evaluation and results","[{\"question\":\"What problem do limited memory language models (LMLMs) address compared to conventional LLMs?\",\"answer\":\"They externalize factual knowledge into a knowledge base (KB) rather than memorizing facts in model weights, then fetch knowledge during generation as needed.\"},{\"question\":\"How does CO-LMLM differ from REL-LMLM in knowledge retrieval?\",\"answer\":\"CO-LMLM uses continuous vector queries and an unstructured index, while REL-LMLM relies on a relational KB queried via explicit decoded natural-language relational tuples.\"},{\"question\":\"Why does CO-LMLM remove restrictions tied to Wikipedia in prior work?\",\"answer\":\"It uses an annotation pipeline that tags free-form factual spans in arbitrary text, eliminating the earlier restriction to Wikipedia-centered relational structures.\"},{\"question\":\"What benefits does CO-LMLM show in experimental evaluation?\",\"answer\":\"Across scales, it improves perplexity and factual precision versus both standard LLMs and prior LMLMs, while preserving downstream NLU performance; results at 360M match SimpleQA-verified quality comparable to strong model baselines.\"}]",1784194480,81,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":91,"head_meta":93,"extra_data":95,"updated_unix":28},"co-lmlm-continuous-query-limited-memory-language-models","",{"@graph":36,"@context":90},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/co-lmlm-continuous-query-limited-memory-language-models/84266/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82,86],{"name":73,"@type":74,"acceptedAnswer":75},"What problem do limited memory language models (LMLMs) address compared to conventional LLMs?","Question",{"text":76,"@type":77},"They externalize factual knowledge into a knowledge base (KB) rather than memorizing facts in model weights, then fetch knowledge during generation as needed.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does CO-LMLM differ from REL-LMLM in knowledge retrieval?",{"text":81,"@type":77},"CO-LMLM uses continuous vector queries and an unstructured index, while REL-LMLM relies on a relational KB queried via explicit decoded natural-language relational tuples.",{"name":83,"@type":74,"acceptedAnswer":84},"Why does CO-LMLM remove restrictions tied to Wikipedia in prior work?",{"text":85,"@type":77},"It uses an annotation pipeline that tags free-form factual spans in arbitrary text, eliminating the earlier restriction to Wikipedia-centered relational structures.",{"name":87,"@type":74,"acceptedAnswer":88},"What benefits does CO-LMLM show in experimental evaluation?",{"text":89,"@type":77},"Across scales, it improves perplexity and factual precision versus both standard LLMs and prior LMLMs, while preserving downstream NLU performance; results at 360M match SimpleQA-verified quality comparable to strong model baselines.","https://schema.org",{"og:url":52,"og:type":92,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":94,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":97},[98,102,106,110,114,119,124,127,132,135,139],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Comic",60,"comic",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},6,"Technology",50,"technology",{"id":120,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":122,"slug":123},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":125,"slug":126},30,"research-report",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":130,"slug":131},9,"Religion & Spirituality",20,"religion-spirituality",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":130,"slug":134},"World Cup","world-cup",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":136,"slug":138},10,"Lifestyle","lifestyle",{"id":140,"doc_module":4,"doc_module_name":46,"category_name":141,"show_sort_weight":20,"slug":142},19,"General","general"]