[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81693-en":3,"doc-seo-81693-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81693,1649267921044,"Ava Thompson","https://us-avatar.wpscdn.com/avatar/1800007509477c92dfb?_k=1782875107921204101",8,"Research & Report","Memory-Managed Long-Context Attention","Memory-managed long-context attention is studied through an explicit bounded memory system with a learned query-independent writer, lifecycle control, query-aware reading, calibrated sparse fallback, and frozen-LLM generation from raw evidence. Track A defines a controlled versioned-variable benchmark where last-mention retrieval is wrong by construction; full lifecycle scores 1.000 on all seeds versus a 0.333 lexical baseline and attains 300/300 generation at 146 tokens. Track B evaluates held-out HotpotQA with distractors, where a bounded two-hop selector with a 32-passage cache and calibrated fallback improves F1 by 5.5–16.6 versus dense retrieval and achieves strong long-context efficiency.","Memory-Managed Long-Context Attention: Bounded Editable Memory with a Hard Lifecycle and Calibrated Sparse Fallback  \nJunyi Zou[zoujunyi@zjydiary. cn](zoujunyi@zjydiary. cn)  \nAvrova Donz  \narXiv :2606 .28876v2 [ cs .CL] 10 Jul 2026  \nJuly 9, 2026  \nAbstract  \nWe study memory-managed long-context attention: explicit bounded memory with a learned query-independent writer, lifecycle control, query-aware reading, calibrated sparse fallback, and frozen-LLM generation from raw evidence. Track A is a controlled versioned-variable task where last-mention retrieval is wrong by construction. Its full lifecycle scores 1.000 on all three seeds versus a 0.333 lexical baseline, and generation reaches 300/300 at 146 prompt tokens, compared with 172/300 for full-context reading at 729 tokens. Track B uses held-out HotpotQA questions and train-derived, answer-excluded distractors at natural and 8.2k-word lengths. A learned twohop selector with a bounded 32-passage cache and fallback beats dense retrieval by 5.5–16.6 F1 and reaches 102–116  \n1 Introduction  \nEfficient long-context modeling has progressed rapidly through linear attention, state-space and recurrent hybrids, and sparse attention [15, 23 , 11 , 27 , 30 , 29 , 9] . These methods reduce the per-token state size or the number of attended positions. Yet a compressed state is not a managed memory. A model that folds information into a fixed matrix must still decide which events deserve persistence, how a newer fact overrides an older one, how a protected fact resists later invalid writes, what to evict under a hard capacity bound, and when to abstain from memory and fall back to retrieval.  \nWe study the hypothesis:  \nState compression and memory management are separate design problems. Long-context models need an explicit lifecycle for writes, overwrites, protection, and eviction, plus a calibrated fallback for content that carries no write-time signal — not only a cheaper attention state.  \nEarlier versions of this preprint reported controlled mechanism evidence but no integrated system: trainable event scoring and hard lifecycle execution existed only as separate experiments, and an opendomain selector had not been demonstrated. This version reports the completed integration under preregistered gates, in two instantiations with frozen 8B/14B backbones: on synthetic lifecycle text (Track A) all five components above execute in one path, and on real multi-hop QA text (Track B) the framework instantiates as a bounded salience cache with learned selection and calibrated fallback (static text exercises no overwrite semantics) .  \nContributions.  \n1. A single implemented path combining a query-independent learned writer, hard bounded lifecycle (overwrite / protection / eviction at 32 slots), query-aware reading, calibrated sparse fallback, and frozen-LLM generation from raw selected evidence — with the full lifecycle exercised on controlled text (Track A) and a bounded-cache instantiation on real text (Track B) (§2) .  \n2. Track A: a versioned variable-tracking benchmark in which naive last-mention retrieval is wrong by construction (alias queries, stale re-mentions, rejected writes, protected slots, capacity pressure), with preregistered gates that the full path passes on every seed while every non-learned and no-lifecycle baseline fails (§3) .  \n3. Track B: held-out HotpotQA questions at two length regimes (extended contexts use trainderived, answer-excluded distractors), where a learned two-hop selector under a 32-passage state bound with calibrated fallback beats budget-matched dense retrieval on every seed of two reader models and, at LongBench-scale length, beats the same model reading the full context at a tenth of the evidence budget; the margins are the selector’s, and the bounded cache is shown topreserve them (§4) .  \n4. A quantified boundary: on static text, write-worthiness without the query is near chance (AUC  \n0.63–0.66 vs. 0.89–0.97 query-aware), so bounded memory alone r","cbCairs1V3YoyZR7","https://ap.wps.com/l/cbCairs1V3YoyZR7","pdf",586006,2,1,17,"English","en",105,"# Abstract\n# Introduction\n# The Implemented Path\n## Components","[{\"question\":\"What is the main idea behind memory-managed long-context attention in this work?\",\"answer\":\"It introduces an explicit bounded memory system with a learned writer and a controlled lifecycle (overwrite, protection, eviction), coupled with query-aware reading and a calibrated sparse fallback when no write-time signal exists.\"},{\"question\":\"How does Track A evaluate the lifecycle mechanism?\",\"answer\":\"Track A uses a versioned-variable benchmark where naive last-mention retrieval is wrong by construction, including alias queries, stale re-mentions, rejected writes, protected slots, and capacity pressure, with preregistered gates validating the full integrated path.\"},{\"question\":\"What approach does Track B use for real multi-hop QA and long contexts?\",\"answer\":\"Track B uses held-out HotpotQA questions with distractors and a learned two-hop selector under a bounded 32-passage cache plus calibrated fallback, which is shown to outperform budget-matched dense retrieval and to reduce evidence budget at long lengths.\"}]",1784175448,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"memory-managed-long-context-attention","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/memory-managed-long-context-attention/81693/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the main idea behind memory-managed long-context attention in this work?","Question",{"text":75,"@type":76},"It introduces an explicit bounded memory system with a learned writer and a controlled lifecycle (overwrite, protection, eviction), coupled with query-aware reading and a calibrated sparse fallback when no write-time signal exists.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does Track A evaluate the lifecycle mechanism?",{"text":80,"@type":76},"Track A uses a versioned-variable benchmark where naive last-mention retrieval is wrong by construction, including alias queries, stale re-mentions, rejected writes, protected slots, and capacity pressure, with preregistered gates validating the full integrated path.",{"name":82,"@type":73,"acceptedAnswer":83},"What approach does Track B use for real multi-hop QA and long contexts?",{"text":84,"@type":76},"Track B uses held-out HotpotQA questions with distractors and a learned two-hop selector under a bounded 32-passage cache plus calibrated fallback, which is shown to outperform budget-matched dense retrieval and to reduce evidence budget at long lengths.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]