[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83992-en":3,"doc-seo-83992-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83992,7971461740886,"Theodore","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","Memory in the Loop In-Process Retrieval as Extended Working Memory for Language Agents","Language agents operate via an observe–reason–act loop, yet their memory is commonly accessed outside the loop as a store queried only once per turn. This work studies moving read/write memory inside the loop at every reasoning step. The primary barrier is latency: networked vector stores can multiply end-to-end delay up to 83×, while per-step overhead collapses when the store is in-process (~100µs). Causal experiments show redundant actions rise with higher store latency, and in-loop memory improves recall across GPT-5-class models.","Memory in the Loop: In-Process Retrieval as Extended Working Memory for Language Agents  \nYusuf Khan  \n[yusuf@mykhan.me](yusuf@mykhan.me)  \nCarlo Lipizzi  \n[clipizzi@stevens.edu](clipizzi@stevens.edu)  \narXiv :2607 .05690v 1 [ cs .AI] 6 Jul 2026  \nAbstract  \nLanguage agents run a loop—observe, reason, act—but the memory they reason over is treated as something outside that loop: a store queried at most once per turn. We study the regime in which memory moves inside the loop, read and written on every reasoning step. The obstacle has always been latency: networked vector stores answer in tens to hundreds of milliseconds, and in-loop retrieval has been shown to inflate end-to-end latency by up to 83 × when retrieval is itself expensive. Prior responses manage that cost: serving-layer scheduling hides it, and“memory-first” designs ration retrieval to once per turn. We argue the cost itself is an assumption rather than a law. Latency is a property of where the store lives, not of the in-loop pattern; an in-process store answers in ∼ 100µs, three orders of magnitude below the network regime, and at that speed the per-step tax collapses.  \nWe ground the distinction in the extended-mind thesis: by the parity principle, an external resource is constitutive of cognition only when it is constantly available, directly accessible without difficulty, and automatically endorsed—criteria whose first two we read as a latency budget. A 100ms store is a tool an agent consults;  \na 100µs store, in a loop wired to consult it, is extended working memory. We then show the premise is causal: holding a fixed per-turn memory-latency budget and varying only the store’s answer speed, redundant actions rise monotonically with store latency—0 .0 of 12 at in-process speed, 7.2 of 12 at a 110ms cloud round trip, where not one lookup fits the budget (gpt-5-nano, gpt-5-mini ; five seeded workloads per rung; exact permutation p=0 .0079; zero guard errors)—and even a 500 ms budget still leaks 1.6 of 12. We demonstrate the regime end-to-end:  \nacross four GPT-5-class models under a bounded context window, recall improves from 0/5 (all forty baseline and window-aware runs) to 3 .6–4.8/5 with in-loop memory, live store ops at p50 80–165µs; an instructed restate-every-reply baseline solves this five-fact task perfectly, which we report and analyze—restatement pays per-turn rent that grows with the working set, exactly the cost the store avoids.  \nThe store never lost a fact in any run (244 of 244 writes kept); every observed miss is a stored fact the agent’s single bounded read never surfaced—a read-policy failure, not a memory failure. Our measurements also relocate the bottleneck: the dominant per-step cost is embedding (∼200–400ms over the network); pairing the in-process store with a small local embedder returns the complete operation to a measured ∼40µs.  \n1 Introduction  \nLanguage agents are defined by a loop: observe, reason, act, repeat. Yet the memory they reason over is typically treated as something outside that loop—a database the agent queries once per turn and otherwise leaves alone. This paper asks what happens when memory moves inside the loop: when an agent can read and write an associative store on every step of its reasoning, as cheaply as it accesses its own context window.  \nPreprint.  \nWe call this regime memory in the loop, echoing—and extending—the familiar “human in the loop.”The obstacle has always been latency. A networked vector store answers in 50–200ms (§2); an agent that consults it at every step pays that cost repeatedly, and recent work shows in-loop retrieval can inflate end-to-end latency by up to 83 × [Yang et al., 2025] . The field has answered on two fronts. Systems work keeps retrieval in the loop and hides its cost at the serving layer: SearchAgent-X [Yang et al., 2025] schedules requests by priority and makes retrieval non-stalling. Industry “memory-first”guidance instead moves memory out of the loop, into a layer queried on","cbCaisF1HqCw3tmU","https://ap.wps.com/l/cbCaisF1HqCw3tmU","pdf",762410,5,1,18,"English","en",105,"# Abstract\n# 1 Introduction\n## Memory in the loop vs outside-the-loop stores\n## Latency as the determining constraint\n## Extended-mind grounding","[{\"question\":\"What does “memory in the loop” mean in this paper?\",\"answer\":\"It means an agent can read from and write to an associative memory store on every reasoning step inside the observe–reason–act loop, rather than querying an external store only once per turn.\"},{\"question\":\"Why does the paper claim latency—not the loop pattern—causes the main obstacle?\",\"answer\":\"Networked stores add significant round-trip delay on repeated in-loop retrievals, inflating end-to-end latency. An in-process store answers in ~100µs, making the per-step retrieval tax negligible and eliminating the need to ration memory access.\"},{\"question\":\"How do the experiments connect store latency to task outcomes?\",\"answer\":\"With a fixed per-turn memory latency budget, varying only the store’s answer speed leads to monotonic increases in redundant actions as latency grows. When latency rises (e.g., ~110ms cloud trips), lookups no longer fit the budget and performance degrades, while in-process speed supports strong end-to-end recall improvements.\"}]",1784191904,45,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"memory-in-the-loop-in-process-retrieval-as-extended-working-memory-for-language-agents","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/memory-in-the-loop-in-process-retrieval-as-extended-working-memory-for-language-agents/83992/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What does “memory in the loop” mean in this paper?","Question",{"text":76,"@type":77},"It means an agent can read from and write to an associative memory store on every reasoning step inside the observe–reason–act loop, rather than querying an external store only once per turn.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"Why does the paper claim latency—not the loop pattern—causes the main obstacle?",{"text":81,"@type":77},"Networked stores add significant round-trip delay on repeated in-loop retrievals, inflating end-to-end latency. An in-process store answers in ~100µs, making the per-step retrieval tax negligible and eliminating the need to ration memory access.",{"name":83,"@type":74,"acceptedAnswer":84},"How do the experiments connect store latency to task outcomes?",{"text":85,"@type":77},"With a fixed per-turn memory latency budget, varying only the store’s answer speed leads to monotonic increases in redundant actions as latency grows. When latency rises (e.g., ~110ms cloud trips), lookups no longer fit the budget and performance degrades, while in-process speed supports strong end-to-end recall improvements.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]