[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85347-en":3,"doc-seo-85347-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85347,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","Extending LLM Context via Associative Recurrent Memory","Extending the context length of large language models (LLMs) is essential for real-world applications, but standard transformers face quadratic compute growth and linear memory scaling that limit long-context processing. This work studies the Associative Recurrent Memory Transformer (ARMT) as a practical mechanism for long-context inference with constant memory scaling and improved efficiency. Three contributions include new long-context datasets, a full ARMT training recipe, and experiments showing better long-length generalization and 30% lower FLOPs.","Extending LLM Context via Associative Recurrent Memory  \nGleb Kuzmin1,4,8 Ivan Rodkin2,6 Aydar Bulatov3,6 Yuri Kuratov3,6 Lyudmila Rvanova1 Mikhail Katkov9,10 Ilia Sochenkov7 Misha Tsodyks9,10 Timothy Baldwin2 Mikhail Burtsev5 Artem Shelmanov2  \n1FusionBrain Lab 2MBZUAI 3 Cognitive AI Systems Lab 4RUDN  \n5London Institute for Mathematical Sciences 6MIRAI 7Lomonosov Moscow State University  \n8Laboratory for Analysis and Controllable Text Generation Technologies RAS  \n9 School of Natural Sciences, Institute for Advanced Study, Princeton  \n10Department of Brain Sciences, Weizmann Institute of Science  \n[kuzmin.gyu@gmail.com](kuzmin.gyu@gmail.com) [artem.shelmanov@mbzuai.ac.ae](artem.shelmanov@mbzuai.ac.ae)  \narXiv :2607 . 1 16 14v 1 [ cs .CL] 13 Jul 2026  \nAbstract  \nExtending the context length of large language models (LLMs) is critical for many real-world applications, yet standard transformers remain constrained by quadratic compute and linear memory scaling. In this work, we investigate the Associative Recurrent Memory Transformer (ARMT) as a practical approach for enabling long-context processing in LLMs, constant memory scaling, and better efficiency.  \nWe make three main contributions. First, we construct two domain-specific long-context datasets designed to evaluate realistic workloads, focusing on narrow-domain fine-tuning scenarios. Second, we propose a comprehensive training recipe for ARMT-based context extension, combining continued pre-training, synthetic long-context data generation, curriculum learning, and selective integration of associative memory into chosen model layers. Third, we present an extensive experimental study demonstrating that ARMT-augmented models: (i) process inputs well beyond their original context limits without degrading performance relative to in-limit baselines; (ii) generalize more effectively to out-of-distribution context lengths; and (iii) need 30% less FLOPs while preserving baseline performance within the original context window.  \n1 Introduction  \nLong-context understanding is crucial for many tasks, such as processing and understanding technical and financial reports, software development, and multi-document reasoning in scientific and legal domains. These scenarios often require models to integrate information distributed across hundreds of thousands or even millions of tokens. However, standard transformer architectures (Vaswani et al., 2017) struggle to scale to such contexts, as the computational and memory costs of self-attention grow quadratically with sequence length. Moreover, transformer performance degrades as the context length increases (Liu et al., 2024 ; Kuratov et al.,  \n2024) . Therefore, since the introduction of the transformer architecture, long-context processing has emerged as a central and rapidly-evolving research direction (Beltagy et al., 2020 ; Katharopoulos et al., 2020 ; Bulatov et al., 2022) . Traditionally, efficient long-context approaches have been built using recurrent architectures (Gu and Dao, 2024 ; Peng et al., 2023); however, such models must typically be trained from scratch, limiting the ability to leverage existing pre-trained LLMs. Moreover, fully-recurrent LMs update the memory at each time step, which complicates high-level information processing in tasks such as structured copying (Jelassi et al., 2024) and instruction following (Park et al., 2024) .  \nRecent studies (Bulatov et al., 2024 ; Rodkinet al., 2024) have explored enhancing transformers with segment-wise context processing and recurrent memory mechanisms. Using human memory as an analogy (Cowan, 2008), full attention within a segment models short-term/working memory, while the module that recurrently propagates crucial information from segment to segment can be viewed as long-term memory. These approaches preserve strong intra-segment modeling performance while enabling linear scaling with respect to context length.  \nIn this work, we focus on the Associative Recurrent Memory ","cbCais3PJqecFM0j","https://ap.wps.com/l/cbCais3PJqecFM0j","pdf",708476,1,33,"English","en",105,"# Introduction\n# Related Work and Background\n# Abstract","[{\"question\":\"Why is long-context processing difficult for standard transformer models?\",\"answer\":\"Self-attention compute and memory usage grow with sequence length, with attention cost scaling quadratically with context size. Performance can also degrade as context length increases.\"},{\"question\":\"What is the Associative Recurrent Memory Transformer (ARMT) used for in this work?\",\"answer\":\"ARMT is used to enable long-context processing by introducing segment-level associative recurrent memory, allowing efficient scaling to extremely long inputs.\"},{\"question\":\"What improvements does the ARMT-augmented approach report in the experiments?\",\"answer\":\"ARMT-augmented models can handle inputs beyond original context limits without degrading in-window performance, generalize better to out-of-distribution context lengths, and achieve about 30% fewer FLOPs while preserving baseline performance within the original window.\"}]",1784202686,83,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"extending-llm-context-via-associative-recurrent-memory","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/extending-llm-context-via-associative-recurrent-memory/85347/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is long-context processing difficult for standard transformer models?","Question",{"text":75,"@type":76},"Self-attention compute and memory usage grow with sequence length, with attention cost scaling quadratically with context size. Performance can also degrade as context length increases.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the Associative Recurrent Memory Transformer (ARMT) used for in this work?",{"text":80,"@type":76},"ARMT is used to enable long-context processing by introducing segment-level associative recurrent memory, allowing efficient scaling to extremely long inputs.",{"name":82,"@type":73,"acceptedAnswer":83},"What improvements does the ARMT-augmented approach report in the experiments?",{"text":84,"@type":76},"ARMT-augmented models can handle inputs beyond original context limits without degrading in-window performance, generalize better to out-of-distribution context lengths, and achieve about 30% fewer FLOPs while preserving baseline performance within the original window.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]