[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86422-en":3,"doc-seo-86422-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},86422,4398048949847,"Eliana","https://ap-avatar.wpscdn.com/avatar/400002536579ef2da7f?_k=1778318612642679267",8,"Research & Report","DiM-WAM World Action Modeling with Diverse Historical Event Memory","World action models (WAMs) predict future visual states and robot actions, yet short local context weakens performance in temporally dependent, long-horizon manipulation. DiM-WAM augments a WAM with Diverse Historical Event Memory (DHEM), using bank-conditioned candidate features, novelty-aware token selection, and accumulated-mass-weighted fusion to keep complementary event tokens within bounded memory. These tokens condition video and action denoising and are shaped by auxiliary supervision for coarse progress cues. Experiments on RMBench and four real-world tasks raise full-task success from 34.8% to 69.8% (and up to 90.0% in real-world evaluation).","DiM-WAM: World Action Modeling with Diverse  \nHistorical Event Memory  \nKai Wang 1 ,2 , Zhaopeng Gu 1 ,2 , Yixiang Chen 1 , Yuan Xu 1 , Qisen Ma 1 , Jiabing Yang 1 , Zhaowen Li2 ,∗ , Yan Huang 1 ,3 ,∗ , Liang Wang 1 , Peng Su2  \narXiv :2606 .27677v2 [ cs .RO] 12 Jul 2026  \nAbstract—World action models (WAMs) jointly predict future visual states and actions, but short local context limits temporally dependent tasks. We introduce DiM-WAM, which augments a WAM with Diverse Historical Event Memory (DHEM). DHEMuses bank-conditioned candidate features, novelty-aware selection, and accumulated-mass-weighted fusion to retain complementary event tokens in bounded memory; these tokens condition video and action denoising, while auxiliary supervision encourages coarse progress cues. In a training-matched comparison with LingBot-VA on RMBench, DiM-WAM improves the average full-task success rate from 34.8% to 69.8%. Under the same demonstration and evaluation protocol on four real-world tasks, it improves the average stage success ratio from 70.6% to 93.5% and the average full-task success rate from 52.5% to 90.0%. Project page: [https://wangkai-casia.github.io/dim-wam/](https://wangkai-casia.github.io/dim-wam/).  \nI. INTRODUCTION  \nVision-language-action models (VLAs) learn robot policies by predicting actions from language instructions and visual observations [1]–[5] . This action-centric formulation enables scalable policy learning, but sparse action labels provide limited supervision for fine-grained manipulation dynamics. WAMs alleviate this limitation by jointly predicting executable actions and future visual states, where future visual prediction provides dense temporal supervision [6]–[8] . However, existing WAMs remain limited in long-horizon tasks, where correct actions often cannot be inferred from current observations or short local contexts alone.  \nConsider the water-dispenser task in Fig. 1. When the robot reaches for the same switch, the observations immediately before the turn-on and turn-off actions can appear visually similar, particularly when transparent water is difficult to perceive. Nevertheless, the required actions differ: the robot should interact with the switch at one stage but avoid or reverse the interaction at another. A WAM relying only on current observations or short contexts may therefore execute incorrect actions or stop at the wrong stage. Resolving this ambiguity requires access to key historical events and task-progress cues. Long-horizon manipulation requires retaining different types of task-relevant events across tasks and stages. For example, some events indicate whether an interaction has occurred or a subtask has been completed, while others preserve information about layouts, target identities, or objectstate changes needed later. Such diversity motivates a memory  \n1Institute of Automation, Chinese Academy of Sciences (CASIA), Beijing, China. 2 Shenzhen Yinwang Intelligent Technology Co., Ltd., Shenzhen, China. 3FiveAges, Beijing, China.  \n∗Corresponding authors: Zhaowen Li and Yan Huang. E-mail: [lizhaowen@yinwang.com](lizhaowen@yinwang.com); [yhuang@nlpr.ia.ac.cn](yhuang@nlpr.ia.ac.cn).  \nFig. 1. Temporal ambiguity in the water-dispenser task: similar local observations near the switch require different actions depending on interaction history and task progress.  \nmechanism that can discover and preserve complementary event patterns.  \nSimply enlarging the local context increases computational cost and may dilute attention to sparse but task-critical events. Memory-augmented VLAs demonstrate the value of history [9]–[12], but action-centric supervision provides limited signals for discovering and organizing heterogeneous historical events. In contrast, WAMs couple action learning with future visual prediction, providing richer supervision over both robot actions and environmental changes. This visual-action signal facilitates learning diverse historical event representations and enables the","cbCaibXk5ZqOJ4YC","https://ap.wps.com/l/cbCaibXk5ZqOJ4YC","pdf",8226012,2,1,"English","en",105,"# Introduction\n# Related Work\n## WAMs","[{\"question\":\"What limitation of existing WAMs motivates DiM-WAM?\",\"answer\":\"Existing WAMs struggle in long-horizon tasks because correct actions may not be determined from current observations or short local contexts alone.\"},{\"question\":\"How does DHEM help resolve temporal ambiguity in manipulation?\",\"answer\":\"DHEM preserves complementary historical event tokens using novelty-aware selection and accumulated-mass-weighted fusion, so task-stage cues and event history can condition video and action denoising.\"},{\"question\":\"What performance improvements does DiM-WAM report?\",\"answer\":\"A training-matched comparison on RMBench increases average full-task success from 34.8% to 69.8%, and real-world evaluations improve average stage success from 70.6% to 93.5% and full-task success from 52.5% to 90.0%.\"}]",1784211662,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"dim-wam-world-action-modeling-with-diverse-historical-event-memory","",{"@graph":35,"@context":84},[36,52,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,46,49],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":20},"https://docshare.wps.com/document/","Document",{"item":47,"name":12,"@type":42,"position":48},"https://docshare.wps.com/document/research-report/",3,{"item":50,"name":13,"@type":42,"position":51},"https://docshare.wps.com/document/dim-wam-world-action-modeling-with-diverse-historical-event-memory/86422/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":23,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What limitation of existing WAMs motivates DiM-WAM?","Question",{"text":74,"@type":75},"Existing WAMs struggle in long-horizon tasks because correct actions may not be determined from current observations or short local contexts alone.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does DHEM help resolve temporal ambiguity in manipulation?",{"text":79,"@type":75},"DHEM preserves complementary historical event tokens using novelty-aware selection and accumulated-mass-weighted fusion, so task-stage cues and event history can condition video and action denoising.",{"name":81,"@type":72,"acceptedAnswer":82},"What performance improvements does DiM-WAM report?",{"text":83,"@type":75},"A training-matched comparison on RMBench increases average full-task success from 34.8% to 69.8%, and real-world evaluations improve average stage success from 70.6% to 93.5% and full-task success from 52.5% to 90.0%.","https://schema.org",{"og:url":50,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":50},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":28,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":28,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]