[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83360-en":3,"doc-seo-83360-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},83360,1099514068365,"Aurelia","https://ap-avatar.wpscdn.com/avatar/10000253d8d9f28188e?_k=1776742907772140068",8,"Research & Report","Temporally Conditioned Memory-Fusion Policies for Visuomotor Learning","Vision–Language–Action policies such as π0.5 and OpenVLA often act reactively, predicting the next action from the current observation, instruction, and proprioceptive state. This fails in stage-dependent manipulation, where visually similar situations require different actions based on latent task progress and past interaction outcomes. The work proposes Temporally Conditioned Memory-Fusion Policies (TFP), using liquid time-constant belief dynamics and event-aware memory updates to modulate action generation. Experiments show higher success rates on LIBERO, LIBERO-plus, and strong performance on memory diagnostics.","TFP: Temporally Conditioned Memory-Fusion Policies for Visuomotor Learning  \nYushen Liang 1,†, Yue Peng 1,†, Baosheng Jin 1,†, Tianluo Zhang 1 , Xinyu Zhang2 , Shuyi Zhou3 , Zhuoran Chen 1 , Xinqi Liu 1 , Shenji Wan 1  \n1NYU Shanghai, Shanghai, China  \n2University of Electronic Science and Technology of China, Chengdu, China  \n3Beijing Institute of Technology, Beijing, China  \n†Equal contribution  \n§ Code: [github.com/Mirage415/TFP-Temporally-conditioned-Memory-Fusion-Policies-for-Visuomotor-Learning](github.com/Mirage415/TFP-Temporally-conditioned-Memory-Fusion-Policies-for-Visuomotor-Learning)  \narXiv :2607 .08283v 1 [ cs .RO] 9 Jul 2026  \nAbstract—Vision–Language–Action (VLA) policies such as π0.5 and OpenVLA perform well on many manipulation tasks, but they are often reactive: the next action is predicted from the current observation, instruction, and proprioceptive state. This assumption breaks down in stage-dependent manipulation, where visually similar states may require different actions depending on latent task progress and previous interaction outcomes. We argue that such tasks require not only memory, but dynamicsaware belief updates: the policy should preserve task progress during stable or occluded phases and revise its belief near contact, release, or subgoal transitions. We introduce Temporally Conditioned Memory-Fusion Policies (TFP), a lightweight memory-action framework for VLA backbones. TFP maintainsan episode-local task-progress belief with Liquid Time-Constant dynamics and injects the updated belief directly into the flowmatching action decoder through adaptive modulation. This lets temporally accumulated context shape the generated action chunk, rather than serving only as passive history context. With a 3.3B-parameter model, TFP improves the average success rate from 96.9% to 98.75% on LIBERO and from 91.4% to 93.77% on LIBERO-plus. On the memory-focused MIKASA ShellGameTouch diagnostic, TFP achieves success up to 75.0%. Mechanistic analyses show that write-gain changes near manipulation events are about 6 × larger than far non-event phases, and hidden-state interventions show that the belief causally modulates generated action chunks. These results suggest that compact, event-sensitive memory dynamics can improve VLA policies under occlusion, visual perturbation, and stage-dependent task structure.  \nIndex Terms—Imitation Learning, Manipulation, MemoryAugmented VLA  \nI. INTRODUCTION  \nVision–Language–Action (VLA) models, such as π0.5 , OpenVLA, and Octo, have achieved strong performance in robotic manipulation by mapping language instructions and multimodal observations directly to actions [2, 3, 8, 16, 22] . However, many VLA policies remain largely reactive: each action is predicted from the current observation, instruction, and proprioceptive state. This is insufficient for stagedependent manipulation, where the same visible scene may require different actions depending on latent task progress and previous interaction outcomes. For example, in object swapping, a visually similar state may correspond to moving the first object to a buffer, moving the second object into the first object’s original location, or terminating the task.  \nThis motivates adding memory, or task belief, to VLA policies. We use “belief” to denote a learned latent memory state that summarizes hidden task progress relevant to future actions. However, robotic manipulation requires more than storing past observations. A useful belief must also decide when to change. During stable transport or occlusion, the policy should preserve task-progress information despite ambiguous visual evidence; near contact, release, or subgoal completion, it should rapidly incorporate new evidence. Thus, the key question is not only whether a VLA has memory, but whether its memory update follows the event structure and temporal dynamics of manipulation.  \nExisting memory-aware robot policies mainly either retrieve from history buffers or maintain recu","cbCaifyaQoQkwkUv","https://ap.wps.com/l/cbCaifyaQoQkwkUv","pdf",2501267,1,15,"English","en",105,"# Introduction\n## Motivation: Reactive VLA Limitations\n## Event-Driven Memory and Belief Updates\n## Dynamics-Aware Belief Tracking with LTC\n## Proposed Framework: TFP\n## Evaluation on Manipulation Benchmarks","[{\"question\":\"Why do reactive VLA policies struggle with stage-dependent manipulation?\",\"answer\":\"In stage-dependent tasks, visually similar states can require different actions depending on latent task progress and previous interaction outcomes, but reactive policies only use the current observation and instruction.\"},{\"question\":\"How does TFP represent and update task progress memory?\",\"answer\":\"TFP maintains an episode-local latent belief with Liquid Time-Constant dynamics, where elapsed time controls how strongly the belief is retained or revised from new observations.\"},{\"question\":\"What role does the updated belief play in action generation?\",\"answer\":\"The updated belief is projected into the action-head conditioning space and injected via adaptive modulation, shaping the flow-matching action distribution so temporally accumulated context actively influences actions.\"}]",1784186989,38,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"temporally-conditioned-memory-fusion-policies-for-visuomotor-learning","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/temporally-conditioned-memory-fusion-policies-for-visuomotor-learning/83360/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do reactive VLA policies struggle with stage-dependent manipulation?","Question",{"text":75,"@type":76},"In stage-dependent tasks, visually similar states can require different actions depending on latent task progress and previous interaction outcomes, but reactive policies only use the current observation and instruction.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does TFP represent and update task progress memory?",{"text":80,"@type":76},"TFP maintains an episode-local latent belief with Liquid Time-Constant dynamics, where elapsed time controls how strongly the belief is retained or revised from new observations.",{"name":82,"@type":73,"acceptedAnswer":83},"What role does the updated belief play in action generation?",{"text":84,"@type":76},"The updated belief is projected into the action-head conditioning space and injected via adaptive modulation, shaping the flow-matching action distribution so temporally accumulated context actively influences actions.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]