[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84940-en":3,"doc-seo-84940-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84940,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","NativeMEM: Native Memory Compression for Long-Horizon Robotic Manipulation","NativeMEM addresses how pretrained Vision-Language-Action (VLA) models preserve long-horizon visual histories with frequent updates while keeping inference efficient. It replaces external memory management with a real-time, long-term memory design that compresses each historical frame-view observation into a single native token using an efficient Native Memory Compression scheme built from the VLA’s own vision encoder. Appended to the input token sequence, memory tokens enable long-term attention with negligible latency overhead. Training uses a two-stage pipeline with a frozen VLA memory tokenizer and subsequent task-specific fine-tuning, achieving major gains in simulation and on real robots while maintaining low GPU memory usage and strong data efficiency.","arXiv :2607 .06678v 1 [ cs .RO] 7 Jul 2026  \nNativeMEM: Native Memory Compression for Long-Horizon Robotic Manipulation  \nZiye Wang1 Modi Shi2 Chaojun Ni3 Jiazhi Yang4 Mengdi Li5 Zhizhong Su5  \nTianwei Lin5 Hongyang Li 1†  \n1 The University of Hong Kong 2 Beihang University 3 Peking University  \n4 The Chinese University of Hong Kong 5 Horizon Robotics  \n[https://opendrivelab.com/NativeMEM](https://opendrivelab.com/NativeMEM)  \nAbstract: How can pretrained Vision-Language-Action (VLA) models retain long-horizon visual histories with high-frequency updates without sacrificing efficiency? Existing approaches rely on external memory management, which restrains either the memory horizon or the reactiveness of pretrained policies. To this end, we present NATIVEMEM, a VLA policy that features long-term and real-time updated memory. At its core is an efficient memory encoding scheme, Native Memory Compression, which repurposes the VLA’s own vision encoder to compress each historical frame from each camera view into a single token. Appended to the input sequence, these memory tokens enable the pretrained VLAto attend over long-term history with negligible latency overhead, requiring neither an external planner nor a freshly initialized memory module. To align the memory tokens with the pretrained policy, we first develop a generic memory tokenizer under the supervision of a frozen VLA on memory-demanding data, and then unfreeze the VLA for task-specific fine-tuning. NATIVEMEM consistently outperforms prior methods, boosting success rates from 32.4% to 84.0% in simulation and up to 98.7% on real robots, while maintaining low inference latency and GPU memory usage. Notably, NATIVEMEM exhibits high data efficiency by achieving competitive results with prior arts using only 20% of the training data.  \nKeywords: VLA Models, Memory Modeling, Long-Horizon Manipulation  \nFigure 1: NATIVEMEM differs from prior memory-augmented VLAs that rely on VLM-generated textual notes or external memory modules. (a) By repurposing the VLA’s own vision encoder, it compresses each historical frame-view observation into a single native memory token, allowing the policy to condition on the full visual history through its original token sequence. (b) This ultracompact representation enables minute-level histories with over 160 frames, providing a 9 × ∼ 40× longer history horizon than prior methods. (c) NATIVEMEM achieves the highest success rates across memory-dependent manipulation tasks.  \n†Corresponding author: Hongyang Li hongyang@hku .hk  \n1 Introduction  \nVision-Language-Action (VLA) models [1, 2, 3, 4, 5, 6, 7, 8, 9, 10] extend Vision-Language Models [11, 12, 13, 14] to embodied decision-making, offering a promising path toward generalist robot control. However, most pretrained VLAs remain reactive, conditioning only on the current observation and instruction [15, 16, 17] . This single-frame setup is insufficient for memory-dependent manipulation, where actions may depend on task progress, prior interactions, counts, occluded states, or failures, requiring action-relevant visual history.  \nTo address this challenge, recent works have explored two main paradigms for memory-augmented VLAs. The first builds memory outside the policy, where a high-level VLM retrieves past keyframesand plans sparse subtasks for a low-level VLA controller [18, 19, 20, 21, 22, 23, 24] . While effective for extending temporal context, such systems often turn memory into text format. However, such subtask descriptions require costly annotations for training, and subtle details are difficult to faithfully encode in language only. Another line of work builds memory with freshly initialized modules inside the policy, through recurrent states [25], compressed histories [26, 9, 27, 28, 29], or retrievalaugmented memory banks [30, 31, 32, 33, 34, 35] . Since these newly introduced modules are unseen during policy pretraining, this paradigm not only increases the overall architectural co","cbCaim8Z2ulpXy66","https://ap.wps.com/l/cbCaim8Z2ulpXy66","pdf",6257358,1,13,"English","en",105,"# Introduction\n## Problem: reactive single-frame conditioning\n## Prior paradigms for memory-augmented VLAs\n## Proposed approach: Native Memory Compression\n## Two-stage training pipeline","[{\"question\":\"What limitation of pretrained VLA models motivates NativeMEM?\",\"answer\":\"Most pretrained VLAs are reactive and condition only on the current observation and instruction, which is insufficient for memory-dependent manipulation requiring task progress and prior visual context.\"},{\"question\":\"How does NativeMEM store and update long-horizon history efficiently?\",\"answer\":\"It compresses each historical frame-view observation into a single native memory token using Native Memory Compression repurposed from the VLA’s own vision encoder, enabling real-time updated memory without external planners or freshly initialized modules.\"},{\"question\":\"How is NativeMEM trained to align memory tokens with the pretrained policy?\",\"answer\":\"It uses a two-stage pipeline: first freeze the VLA and train a generic memory tokenizer under supervision, then unfreeze the VLA for task-specific fine-tuning so the tokens align with the action prediction objective.\"}]",1784199593,33,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"nativemem-native-memory-compression-for-long-horizon-robotic-manipulation","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/nativemem-native-memory-compression-for-long-horizon-robotic-manipulation/84940/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What limitation of pretrained VLA models motivates NativeMEM?","Question",{"text":75,"@type":76},"Most pretrained VLAs are reactive and condition only on the current observation and instruction, which is insufficient for memory-dependent manipulation requiring task progress and prior visual context.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does NativeMEM store and update long-horizon history efficiently?",{"text":80,"@type":76},"It compresses each historical frame-view observation into a single native memory token using Native Memory Compression repurposed from the VLA’s own vision encoder, enabling real-time updated memory without external planners or freshly initialized modules.",{"name":82,"@type":73,"acceptedAnswer":83},"How is NativeMEM trained to align memory tokens with the pretrained policy?",{"text":84,"@type":76},"It uses a two-stage pipeline: first freeze the VLA and train a generic memory tokenizer under supervision, then unfreeze the VLA for task-specific fine-tuning so the tokens align with the action prediction objective.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]