[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81621-en":3,"doc-seo-81621-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81621,34359740700684,"Finn","https://ap-avatar.wpscdn.com/avatar/1f400023980c374ae676?_k=1777273430885731487",8,"Research & Report","RoboStream Weaving Spatio-Temporal Reasoning with Memory in Vision-Language Models for Robotics","Enabling reliable long-horizon robotic manipulation for open-world embodied intelligence requires persistent spatial and causal awareness. Vision-language model planners often treat each step as an isolated observation-to-action mapping, repeatedly reinferring geometry from raw pixels and lacking memory of how prior actions reshape the environment. As a result, perceptual errors accumulate, temporarily occluded objects are forgotten, and compounding failures cascade into precondition violations. RoboStream introduces STF-Tokens for geometric anchoring and a Causal Spatio-Temporal Graph for action-triggered state transitions, enabling object permanence under occlusion without extra training.","arXiv :2603 . 12939v2 [ cs .RO] 10 Jul 2026  \nRoboStream: Weaving Spatio-Temporal Reasoning with Memory in Vision-Language Models for  \nRobotics  \nYuzhi Huang 1 ,2∗‡♠ Jie Wu 1 ,2∗‡ Weijue Bu3∗ Ziyi Xiong2 ,4‡ Gaoyang Jiang5 Ye Li 1 Kangye Ji 1 ,2‡ Shuzhao Xie 1 Yue Huang6 Chenglei Wu2 Jingyan Jiang4† Zhi Wang 1†  \n1 Shenzhen International Graduate School, Tsinghua University  \n2 YuanxingGuangnian Robotics  \n3 China University of Mining and Technology  \n4 Shenzhen Technology University  \n5 Huazhong University of Science and Technology  \n6 Xiamen University  \nWebsite: [https://robostream123.github.io/](https://robostream123.github.io/)  \nAbstract. Enabling reliable long-horizon robotic manipulation is a crucial step toward open-world embodied intelligence. However, VLM-based planners treat each step as an isolated observation-to-action mapping, forcing them to reinfer scene geometry from raw pixels at every decision step while remaining unaware of how prior actions have reshaped the environment. Despite strong short-horizon performance, these systems lack the spatio-temporal reasoning required for persistent geometric anchoring and memory of action-triggered state transitions. Without persistent state tracking, perceptual errors accumulate across the execution horizon, temporarily occluded objects are catastrophically forgotten, and compounding failures lead to precondition violations that cascade through subsequent steps. In contrast, humans maintain a persistent mental model that continuously tracks spatial relations and action consequences across interactions rather than reconstructing them at each instant. Inspired by this human capacity for causal spatio-temporal reasoning with persistent memory, we propose RoboStream, a training-free framework that achieves geometric anchoring through Spatio-Temporal Fusion Tokens (STF-Tokens), which bind visual evidence to 3D geometric attributes for persistent object grounding, and maintains causal continuity via a Causal Spatio-Temporal Graph (CSTG) that records action-triggered state transitions across steps. This design enables the planner to trace causal chains and preserve object permanence under occlusion without additional training or fine-tuning. RoboStream achieves a 90 .5% success rate on long-horizon RLBench tasks and a 44 .4% success rate on challenging real-world block-building tasks, where both SoFar and VoxPoser score 11.1%, demonstrating that spatio-temporal reasoning and causal memory are critical missing components for reliable long-horizon manipulation.  \nKeywords: Robot Manipulation · Vision-Language Models · Long  \nHorizon Planning · Spatio-Temporal Reasoning · Causal Memory  \n∗ Equal contribution. † Corresponding author. ♠ Project lead.‡ Work done during an internship at Yuanxing Robotics.  \n2 Y. Huang et al.  \n1 Introduction  \nVision-language models (VLMs) have emerged as powerful planners for robotic manipulation, leveraging internet-scale semantic knowledge and visual perception to translate high-level instructions into executable action sequences [14, 32 , 34 , 81] . VLM-based systems have demonstrated strong performance on short-horizon manipulation tasks, including precise grasping, pose-aware placement, and decomposition of language instructions into subgoal sequences [4, 26 , 47 , 56] .  \nHowever, this success does not readily extend to long-horizon tasks, which demand sustained spatial and causal awareness across multi-step action sequences that irreversibly alter the environment [55, 66 , 76] . Current VLM-based planners lack the persistent state tracking required to bridge observations across steps, forcing them to reinfer world state from partial observations at each decision step [19, 28 , 79] . Without accumulated context, perceptual errors and action effects compound across steps, ultimately resulting in cascading failures, as illustrated in Fig. 1. In contrast, humans maintain a coherent mental model [35] that continuously integrates spatial relation","cbCaidi3HsOCXEIH","https://ap.wps.com/l/cbCaidi3HsOCXEIH","pdf",32104670,5,1,41,"English","en",105,"# Introduction\n## Problem: Limitations of step-wise VLM planning\n## Proposed approach: RoboStream with STF-Tokens and CSTG","[{\"question\":\"Why do VLM-based planners struggle with long-horizon robotic manipulation?\",\"answer\":\"They often operate as step-wise observation-to-action mappings, reinferring geometry from raw pixels each decision step and lacking persistent tracking of how actions change the environment. This causes accumulated perceptual errors and cascading precondition violations.\"},{\"question\":\"What is RoboStream, and how does it enable persistent geometric grounding?\",\"answer\":\"RoboStream is a training-free framework that uses Spatio-Temporal Fusion Tokens (STF-Tokens) to bind visual evidence to 3D geometric attributes, producing identity-persistent object grounding across steps.\"},{\"question\":\"How does RoboStream maintain causal continuity across action sequences?\",\"answer\":\"It maintains a Causal Spatio-Temporal Graph (CSTG) that records action-triggered state transitions across steps, allowing the planner to trace causal chains and preserve object permanence even under occlusion.\"}]",1784174881,103,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"robostream-weaving-spatio-temporal-reasoning-with-memory-in-vision-language-models-for-robotics","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/robostream-weaving-spatio-temporal-reasoning-with-memory-in-vision-language-models-for-robotics/81621/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why do VLM-based planners struggle with long-horizon robotic manipulation?","Question",{"text":76,"@type":77},"They often operate as step-wise observation-to-action mappings, reinferring geometry from raw pixels each decision step and lacking persistent tracking of how actions change the environment. This causes accumulated perceptual errors and cascading precondition violations.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What is RoboStream, and how does it enable persistent geometric grounding?",{"text":81,"@type":77},"RoboStream is a training-free framework that uses Spatio-Temporal Fusion Tokens (STF-Tokens) to bind visual evidence to 3D geometric attributes, producing identity-persistent object grounding across steps.",{"name":83,"@type":74,"acceptedAnswer":84},"How does RoboStream maintain causal continuity across action sequences?",{"text":85,"@type":77},"It maintains a Causal Spatio-Temporal Graph (CSTG) that records action-triggered state transitions across steps, allowing the planner to trace causal chains and preserve object permanence even under occlusion.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]