[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83866-en":3,"doc-seo-83866-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83866,8796095462418,"Noah","https://ap-avatar.wpscdn.com/avatar/80000253c1241d02b47?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778826106357471780",8,"Research & Report","DSWAM: A Dual-System World Action Foundation Model for Fine-Grained Robot Manipulation","World Action Models (WAMs) enable video-based world modeling as dense supervision for robot action learning, but often miss the explicit language-level planning interface needed to decompose coarse instructions into fine-grained subtasks. DSWAM introduces a dual-system design: a System 1 WAM executor as the default control path, with an optional System 2 vision-language subtask planner activated only when decomposition is beneficial. The planner predicts executable subtasks from short-term visual history and a global task prompt, while the executor performs world-aware action generation. Real-robot experiments on DeMaVLA show improved folding success and reduced completion time, with System 2 supervision enhancing stability on decomposition-friendly tasks.","arXiv :2607 .04927v 1 [ cs .RO] 6 Jul 2026  \nDSWAM: A Dual-System World Action Foundation Model for Fine-Grained Robot Manipulation  \nJian Zhu 1 ,∗ ,†,‡, Jianjun Zhang 1 ,2 ,∗ , Taiyi Su 1 ,∗ , Tianbin Liu 1 ,∗ , Zhangyuan Wang 1 , Kai Xie 1 , Zitai Huang 1 ,2 , Chong Ma 1 ,2 , Youzhang He 1 , Tianjian Wang 1 , Hanyang  \nWang 1 , Weihao Ding 1 , Yi Xu 1 ,†  \n1 AIRC, Midea Group, 2 Tongji University  \n∗ Equal Contribution, †Corresponding Author, ‡Project Leader,  \nAbstract  \nWorld Action Models (WAMs) provide a promising alternative to Vision-Language-Action (VLA) policies by using video-based world modeling as dense supervision for robot action learning. Existing WAMs excel at physically grounded execution, but typically lack the explicit language-level planning interface in VLM-based VLAs for decomposing coarse instructions. Such decomposition becomes important when household tasks involve complex multi-step goals, where coarse user commands need to be converted into sequences of fine-grained executable subtasks. Meanwhile, the field still lacks a fair real-robot comparison between VLA and WAM execution capabilities, since existing systems often differ in data, robot embodiments, and task protocols. To address both the decomposition gap and the need for a controlled WAM-VLA comparison, we introduce DSWAM, a Dual-System World Action Foundation Model for fine-grained robot manipulation. DSWAM keeps a System 1 WAM executor as the default control path and optionally activatesa System 2 vision-language subtask planner only when task decomposition is useful. The planner predicts executable subtasks from short-term visual history and a global task prompt, while the WAM executor performs world-aware action generation for each instruction or subtask. The executor is trained with action prediction and video co-training, but inference directly predicts action chunks without explicit future video generation. To make this execution path practical on real robots, we further integrate TensorRT acceleration, asynchronous execution, and real-time chunking (RTC) so that policy queries do not block robot control. To provide a fair real-robot comparison with VLA policies, we build and evaluate DSWAM under the DeMaVLA real-world deformable manipulation setting with matched robot platform, pretraining data, post-training data, and evaluation criteria. In this folding benchmark, DSWAM runs in WAM-only mode with the System 2 planner disabled. Under this setting, DSWAM improves average real-world folding success rate from 92 .5% to 96 .3% and reduces completion time from 2′ 18′′ to 1′ 44′′. We further show that the optional System 2 subtask supervision improves the real-world execution stability on tasks that benefit from decomposition by increasing success rate and reducing rollout mistakes.  \nKeywords: World Action Model, Dual-System, Fine-Grained Robot Manipulation, Video Cotraining, TensorRT Acceleration  \nDate: July 7, 2026  \nCorrespondence: Jian Zhu, Yi Xu  \nProject Page: [https://ds-wam.github.io/](https://ds-wam.github.io/)  \n1 Introduction  \nWorld Action Models (WAMs) have recently emerged as a promising alternative to Vision-Language-Action (VLA) policies for robotic manipulation [3, 17 , 18 , 26 , 27] . Existing VLA policies convert visual observations and language instructions into executable robot actions [4, 5 , 9 , 13 , 16 , 21], but their common single-frame, observation-to-action formulation provides limited temporal context for modeling scene evolution during robot manipulation. WAMs address this limitation by learning how visual states evolve under actions, providing world-aware representations of physical dynamics. This video-based formulation gives dense supervision for robot action learning and enables strong performance in contact-rich manipulation tasks such as grasping, placing, and deformable object manipulation, where accurate modeling of physical interactions is critical.  \nHowever, when household tasks involve complex mul","cbCaisMNmMPzts6m","https://ap.wps.com/l/cbCaisMNmMPzts6m","pdf",3113232,5,1,13,"English","en",105,"# Introduction\n## Background: WAM vs VLA\n## Modeling gap: instruction decomposition\n## Real-robot comparison challenges\n## Proposed approach: DSWAM dual-system design","[{\"question\":\"What problem does DSWAM address in existing WAM and VLA approaches?\",\"answer\":\"Existing WAMs focus on execution but lack language-level planning for decomposing coarse instructions into fine-grained subtasks, which is needed for complex multi-step household goals. In addition, prior comparisons between VLA and WAM execution capabilities are often confounded by differing experimental setups.\"},{\"question\":\"How does DSWAM decide when to use the System 2 planner?\",\"answer\":\"DSWAM keeps a System 1 WAM executor as the default control path and activates the System 2 vision-language subtask planner only when task decomposition is useful. For atomic instructions or tasks reliably solved by the executor, it runs in WAM-only mode with the planner disabled.\"},{\"question\":\"What performance gains were reported in the DeMaVLA real-world deformable manipulation benchmark?\",\"answer\":\"In the folding benchmark, DSWAM (WAM-only mode with System 2 disabled) increases average real-world folding success rate from 92.5% to 96.3% and reduces completion time from 2′18″ to 1′44″. Optional System 2 supervision further improves stability by increasing success rate and reducing rollout mistakes on decomposition-benefiting tasks.\"}]",1784191077,33,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"dswam-a-dual-system-world-action-foundation-model-for-fine-grained-robot-manipulation","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/dswam-a-dual-system-world-action-foundation-model-for-fine-grained-robot-manipulation/83866/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does DSWAM address in existing WAM and VLA approaches?","Question",{"text":76,"@type":77},"Existing WAMs focus on execution but lack language-level planning for decomposing coarse instructions into fine-grained subtasks, which is needed for complex multi-step household goals. In addition, prior comparisons between VLA and WAM execution capabilities are often confounded by differing experimental setups.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does DSWAM decide when to use the System 2 planner?",{"text":81,"@type":77},"DSWAM keeps a System 1 WAM executor as the default control path and activates the System 2 vision-language subtask planner only when task decomposition is useful. For atomic instructions or tasks reliably solved by the executor, it runs in WAM-only mode with the planner disabled.",{"name":83,"@type":74,"acceptedAnswer":84},"What performance gains were reported in the DeMaVLA real-world deformable manipulation benchmark?",{"text":85,"@type":77},"In the folding benchmark, DSWAM (WAM-only mode with System 2 disabled) increases average real-world folding success rate from 92.5% to 96.3% and reduces completion time from 2′18″ to 1′44″. Optional System 2 supervision further improves stability by increasing success rate and reducing rollout mistakes on decomposition-benefiting tasks.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]