[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84837-en":3,"doc-seo-84837-105":30,"detail-sidebar-cat-0-en-105":95},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84837,8796095462418,"Noah","https://ap-avatar.wpscdn.com/avatar/80000253c1241d02b47?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778826106357471780",8,"Research & Report","Learning 4D Geometric Priors for Inference-Efficient World Action Models","World Action Models (WAMs) enable robotic manipulation by modeling visual future dynamics together with executable action sequences, yet common co-training largely targets appearance-focused video latents and cannot adequately capture temporally evolving geometry needed for accurate grasping and contact. MECo-WAM introduces action-relevant 4D geometric priors via multi-expert co-training, using a lightweight 4D expert supervised by relational targets from a frozen VGGT encoder and decayed guidance to avoid non-causal shortcuts. Action-aware temporal geometric distillation aligns in-frame relations and their temporal evolution. Deployment removes all auxiliary 4D modules, improving LIBERO and RoboTwin performance without increasing inference cost.","Learning 4D Geometric Priors for Inference-Efficient World Action Models  \nJianjun Zhang1,2,* Jian Zhu2,‡ Taiyi Su2 Chong Ma1,2 Zitai Huang1,2  \nYi Xu2,† Hanli Wang1,†  \n1Tongji University 2AIRC, Midea Group  \n†Corresponding Author ‡Project Leader  \n[Project Page: meco-wam.github.io](Project Page: meco-wam.github.io)  \narXiv :2607 .05468v 1 [ cs .RO] 6 Jul 2026  \nAbstract  \nWorld Action Models (WAMs) have shown strong potential for robotic manipulation by jointly modeling visual future dynamics and executable action sequences. However, existing video-action co-training methods primarily optimize appearance-oriented video latents, which may insufficiently capture the temporally evolving geometry required for precise manipulation. We propose MECo-WAM, a Multi-Expert CoTraining World Action Model that injects action-relevant 4D geometric priors into video-action representations while preserving the original lightweight inference graph. During training, MECo-WAM combines video and action experts with a lightweight 4D expert supervised by relational targets from a frozen VGGT encoder. Asymmetric expert visibility prevents non-causal shortcuts from auxiliary geometry to action generation. To transfer geometric knowledge into the deployed video-action pathway, we introduce decayed 4D read-mask attention, which provides restricted current-frame geometric guidance early in training and progressively removes this dependency. We further propose action-aware temporal geometric distillation, which aligns within-frame geometric relations and their temporal evolution while emphasizing visual regions most relevant to robot actions. At deployment, all auxiliary 4D components are removed. Experiments on LIBERO (98.2%), RoboTwin 2.0 (92.6%), and challenging real-world manipulation tasks show that MECo-WAM improves manipulation performance without increasing inference cost.  \nIntroduction  \nRobotic manipulation requires a policy to map visual observations and language instructions to precise action trajectories (Hu et al. 2025; Ma et al. 2026; Su et al. 2026; Wang et al. 2026; Su et al. 2025) . World action models offer a promising formulation by jointly learning how visual states evolve under interaction and how robot actions should be generated (Ye et al. 2026c; Kim et al. 2026; Bi et al. 2026; Li et al. 2026b) . Compared with direct action policies, video-action co-training can provide richer motion and interaction priors, enabling the policy to reason over changes that unfold beyond a single observation (Ye et al. 2026a; Yuan et al. 2026) .  \nDespite this advantage, the visual representation learned by many WAMs remains dominated by appearance-oriented  \n* This work was completed during an internship at Midea AIResearch Centers.  \n95  \n90  \n80  \n65  \n60  \n150 200 250 1K 2K  \nSuccess Rate (%)  \n\\\\\\\\  \n\\\\  \nInference Latency (ms)  \nFigure 1: Comparison of MECo-WAM with Baselines inaction-chunk inference latency and task success rate on RoboTwin.  \nvideo prediction. A video latent can support plausible future synthesis without explicitly preserving the spatial relations that determine whether a grasp is reachable, whether an object is aligned with a target, or whether contact will cause astable transition. These relations are not static: they evolve asthe robot approaches, contacts, moves, and releases objects. Recent geometry-aware VLA andWAM studies therefore introduce 3D or 4D structure to strengthen spatial grounding for manipulation (Qu et al. 2025; Li et al. 2025a,b; Guo et al. 2026; Li et al. 2026c) . Thus, manipulation-oriented WAMs require geometry-aware temporal representations beyond visual plausibility.  \nA direct approach is to introduce explicit 4D reconstruction or dense geometric prediction into the world action modeling pipeline (Guo et al. 2026; Li et al. 2026c) . However, making geometry an explicit deployment-time output increases inference cost and may shift optimization toward geometric reconstruction that is only weakly co","cbCaikZn2rmdAbWe","https://ap.wps.com/l/cbCaikZn2rmdAbWe","pdf",3095505,2,1,9,"English","en",105,"# Abstract\n# Introduction\n## Problem: appearance-dominated video prediction\n## Approach: MECo-WAM multi-expert co-training\n## Decayed 4D read-mask attention and deployment strategy\n## Action-aware temporal geometric distillation\n## Contributions and experimental evaluation","[{\"question\":\"What limitation in existing WAM co-training motivates MECo-WAM?\",\"answer\":\"Existing methods mainly optimize appearance-oriented video latents, which do not sufficiently represent temporally evolving geometry required for precise robotic manipulation.\"},{\"question\":\"How does MECo-WAM inject 4D geometric priors during training while keeping inference efficient?\",\"answer\":\"MECo-WAM uses a lightweight 4D expert only during training, supervised by relational targets from a frozen VGGT encoder, and restricts early cross-expert access so only limited current-frame geometry guidance is transferred. At deployment, all auxiliary 4D components are removed.\"},{\"question\":\"What is decayed 4D read-mask attention used for?\",\"answer\":\"It provides restricted current-frame geometric guidance early in training and progressively removes that dependency before deployment, preventing non-causal shortcuts.\"},{\"question\":\"How does action-aware temporal geometric distillation improve the learned representations?\",\"answer\":\"It aligns predicted 4D keyframe relations and their temporal evolution while emphasizing visual regions most relevant to robot actions.\"}]",1784198626,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":90,"head_meta":92,"extra_data":94,"updated_unix":28},"learning-4d-geometric-priors-for-inference-efficient-world-action-models","",{"@graph":36,"@context":89},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/learning-4d-geometric-priors-for-inference-efficient-world-action-models/84837/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81,85],{"name":72,"@type":73,"acceptedAnswer":74},"What limitation in existing WAM co-training motivates MECo-WAM?","Question",{"text":75,"@type":76},"Existing methods mainly optimize appearance-oriented video latents, which do not sufficiently represent temporally evolving geometry required for precise robotic manipulation.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does MECo-WAM inject 4D geometric priors during training while keeping inference efficient?",{"text":80,"@type":76},"MECo-WAM uses a lightweight 4D expert only during training, supervised by relational targets from a frozen VGGT encoder, and restricts early cross-expert access so only limited current-frame geometry guidance is transferred. At deployment, all auxiliary 4D components are removed.",{"name":82,"@type":73,"acceptedAnswer":83},"What is decayed 4D read-mask attention used for?",{"text":84,"@type":76},"It provides restricted current-frame geometric guidance early in training and progressively removes that dependency before deployment, preventing non-causal shortcuts.",{"name":86,"@type":73,"acceptedAnswer":87},"How does action-aware temporal geometric distillation improve the learned representations?",{"text":88,"@type":76},"It aligns predicted 4D keyframe relations and their temporal evolution while emphasizing visual regions most relevant to robot actions.","https://schema.org",{"og:url":51,"og:type":91,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":93,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":96},[97,101,105,109,114,119,124,127,131,134,138],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},"Exam",70,"exam",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},5,"Comic",60,"comic",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},6,"Technology",50,"technology",{"id":120,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":122,"slug":123},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":125,"slug":126},30,"research-report",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":129,"slug":130},"Religion & Spirituality",20,"religion-spirituality",{"id":129,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":129,"slug":133},"World Cup","world-cup",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":135,"slug":137},10,"Lifestyle","lifestyle",{"id":139,"doc_module":4,"doc_module_name":46,"category_name":140,"show_sort_weight":110,"slug":141},19,"General","general"]