[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83276-en":3,"doc-seo-83276-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},83276,13056703019662,"Evangeline","https://ap-avatar.wpscdn.com/avatar/be000253a8e92610077?_k=1778726343310543188",8,"Research & Report","LaMem-VLA: Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation","Mainstream Vision-Language-Action (VLA) models infer actions from the current observation under a Markovian assumption, which weakens performance on long-horizon tasks with temporal dependencies. Existing memory-enhanced approaches either enlarge observation windows or retrieve past context as external auxiliary policy signals, leaving memory outside the native latent reasoning space. LaMem-VLA reconstructs historical experience as latent memory tokens and interweaves them directly with VLA reasoning. It organizes history into short- and long-term vaults, retrieves evidence, condenses it into compact tokens, and weaves them into a continuous embedding sequence to guide action generation with bounded context. Experiments on SimplerEnv and LIBERO validate superiority.","arXiv :2607 .07608v 1 [ cs .RO] 8 Jul 2026  \nLAMEM-VLA: DUAL LATENT MEMORY IN VISIONLANGUAGE-ACTION MODELS FOR ROBOTIC MANIPULATION  \nHongyu Qu 1 , Jianzhe Gao2 , Xiaobin Hu3 , Shaohuan Yang 1 , Xinlei Yu3 , Rui Yan 1 , Wenguan Wang2 , Xiangbo Shu 1 , Shuicheng Yan3  \n1Nanjing University of Science and Technology 2Zhejiang University  \n3 National University of Singapore  \nABSTRACT  \nMainstream Vision-Language-Action (VLA) models predict actions primarily from the current observation under a Markovian assumption, thus struggling with long-horizon, temporally dependent tasks. Existing memory-augmented VLAs either expand the observation window or retrieve history from the memory bank as auxiliary policy-side context. However, they leave memory outside the native latent embedding space of VLA reasoning, preventing historical experience from being fluidly interleaved with multimodal reasoning and action formation. To this end, we introduce LaMem-VLA, a latent-memory-native framework that reconstructs historical experience into latent memory tokens and directly interweaves them with VLA reasoning. At its core, LaMem-VLA introduces four coordinated components: (i) a curator that organizes historical experience into two complementary short-term and long-term memory vaults; (ii) a seeker that queries both vaults using the multimodal cognition to retrieve context-relevant evidence; (iii) a condenser that reconstructs the retrieved evidence into compact short-term and long-term latent memory tokens; and (iv) a weaver that injects these memory tokens with the current observation and instruction into one continuous embedding sequence. By representing, retrieving, and consuming historical experience entirely in the same continuous latent space, LaMem-VLA enables memory to directly participate in VLA reasoning and guide action generation under a bounded context. Extensive experiments on SimplerEnv and LIBERO demonstrate the superiority of our LaMem-VLA. The project page will be available at LaMem-VLA.  \n1 INTRODUCTION  \nVision-language-action (VLA) models [1, 2, 3, 4] have become a promising paradigm for general robotic manipulation. By combining the powerful capabilities of pretrained vision-language models [5, 6, 7] with policy learning [8, 9, 10] on robotic data [11, 12, 13, 14], they map visual observations and language instructions into executable action chunks. Despite this progress, most existing VLA models [2, 1, 15] implicitly rely on a Markovian assumption, predicting actions primarily from the current observation without considering temporal dependencies. This simplification creates a temporal short-horizon bias:  \nFigure 1: Paradigm comparison of memory-augmented VLA Models. (a) Unlike previous VLA models that store historical experience in an auxiliary memory bank and consume retrieved memory as external policy-side context,(b) LaMem-VLA treats historical experience as context-native latent memory, which is stored, retrieved, and consumed in the model embedding space.  \nVLA models can react to the currently visible state, but fail to reason about previous state  \ntransitions, completed operation steps, and the current phase of a multi-step task. As a result, these models especially struggle with long-horizon manipulation tasks.  \nRecent efforts have sought to alleviate temporal short-horizon bias by augmenting VLA models [8, 16, 17] with historical context or memory mechanisms along two main axes. (i) One line of work incorporates short-horizon episode context by concatenating historical frames [18, 19] or extending the input into a video sequence [20, 21, 22] . Although such designs expose recent state changes, they incur computational and memory costs that grow with the context length, while the fixed temporal horizon imposes an inherent memory ceiling, causing potentially task-relevant evidence outside the window to be discarded. (ii) Another line of work [23, 24, 25] retrieves past trajectories or relevant historical tokens ","cbCaikq9hziV6m0r","https://ap.wps.com/l/cbCaikq9hziV6m0r","pdf",1124508,1,15,"English","en",105,"# Introduction\n## Memory-augmented VLA and temporal short-horizon bias\n## LaMem-VLA: context-native latent memory tokens\n# Method Overview\n## Curator, seeker, condenser, and weaver components","[{\"question\":\"Why do mainstream VLA models struggle with long-horizon robotic manipulation?\",\"answer\":\"They largely follow a Markovian assumption and predict actions mainly from the current observation, so they do not properly reason over temporal dependencies and earlier state transitions across multi-step tasks.\"},{\"question\":\"What limitation do existing memory-augmented VLA approaches have?\",\"answer\":\"They tend to store historical information outside the model’s native latent embedding space and consume retrieved memory as external auxiliary policy-side context, which prevents memory from being smoothly interleaved with multimodal reasoning.\"},{\"question\":\"How does LaMem-VLA enable historical experience to participate in VLA reasoning?\",\"answer\":\"It reconstructs historical experience into latent memory tokens, retrieves relevant evidence from coordinated short- and long-term vaults, condenses it into compact tokens, and weaves those tokens with the current observation and instruction into one continuous embedding sequence.\"}]",1784186444,38,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"lamem-vla-dual-latent-memory-in-vision-language-action-models-for-robotic-manipulation","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/lamem-vla-dual-latent-memory-in-vision-language-action-models-for-robotic-manipulation/83276/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do mainstream VLA models struggle with long-horizon robotic manipulation?","Question",{"text":75,"@type":76},"They largely follow a Markovian assumption and predict actions mainly from the current observation, so they do not properly reason over temporal dependencies and earlier state transitions across multi-step tasks.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What limitation do existing memory-augmented VLA approaches have?",{"text":80,"@type":76},"They tend to store historical information outside the model’s native latent embedding space and consume retrieved memory as external auxiliary policy-side context, which prevents memory from being smoothly interleaved with multimodal reasoning.",{"name":82,"@type":73,"acceptedAnswer":83},"How does LaMem-VLA enable historical experience to participate in VLA reasoning?",{"text":84,"@type":76},"It reconstructs historical experience into latent memory tokens, retrieves relevant evidence from coordinated short- and long-term vaults, condenses it into compact tokens, and weaves those tokens with the current observation and instruction into one continuous embedding sequence.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]