[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83841-en":3,"doc-seo-83841-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83841,2336464648322,"Aria","https://ap-avatar.wpscdn.com/avatar/2200025388227c56fec?_k=1778556882303663488",8,"Research & Report","KAM-WM: Kinematic Affordance Maps from Latent World Models for Robot Manipulation","Learning manipulation from few demonstrations requires visual priors that indicate not only where interaction should occur, but also how it should begin. KAM-WM extracts a coarse directional cue from a frozen latent video world model by querying a Flow Matching image-to-video backbone once and interpreting its single-step latent velocity as a Kinematic Affordance Map. A lightweight Perceiver compresses these cues into tokens that condition a diffusion policy with RGB and proprioception, yielding strong results on LIBERO and RoboTwin 2.0.","arXiv :2607 .04652v 1 [ cs .RO] 6 Jul 2026  \nKAM-WM: Kinematic Affordance Maps from Latent World Models for Robot Manipulation  \nXinyu Shao 1 ,2 , Keru Zhou 1 , Guowei Huang2 , Yajun Gao2 , Tongtong Cao2 , Xiu Li 1  \n1Tsinghua Shenzhen International Graduate School  \n2Huawei Technologies Ltd.  \n[shaoxy23@mails.tsinghua.edu.cn](shaoxy23@mails.tsinghua.edu.cn)  \nAbstract: Learning manipulation from few demonstrations requires visual priors that capture not only where to interact, but also how the interaction should begin;  \nstatic priors such as segmentation masks encode only the former. We present KAM-WM, a framework that extracts a coarse directional interaction cue from a frozen latent video world model without rollout or world-model fine-tuning. KAMWM queries a Flow Matching image-to-video backbone once and interprets its single-step latent velocity as a Kinematic Affordance Map (KAM), which provides task-conditioned interaction regions and coarse motion structure. A lightweight Perceiver compresses KAM into tokens that condition a diffusion policy together with RGB observations and proprioception. Across LIBERO and RoboTwin 2.0, KAM-WM reaches 90.6% average success on LIBERO and achieves 65.7% and 22.4% success rates in the Easy and Hard settings on RoboTwin 2.0, respectively.  \nControlled comparisons against a zero-order mask prior suggest that part of the gains comes from directional information beyond spatial localization alone. These results indicate that, in the evaluated settings, a frozen video model can provide a useful first-order visual prior for control without the test-time cost of future rollout.  \nKeywords: Robot Manipulation, Latent World Models, Imitation Learning  \n1 Introduction  \nLearning manipulation from few demonstrations requires visual priors that indicate not only where to attend, but also how the interaction should begin. Static priors such as segmentation masks help localization, yet remain zero-order cues: they mark relevant objects or regions without encoding approach direction. This is limiting for tasks such as hanging a mug on a hook, where location alone does not determine a successful motion. As illustrated in Figure 1(a), a bottle mask may cover the whole object, whereas a more useful prior would emphasize the lower graspable region together with the gripper’s approach motion.  \nPretrained video models are an appealing source of this directional information, since future-frame prediction requires encoding how objects move and interact. Most robotics methods that use this knowledge, however, roll out future frames as visual subgoals or co-train the backbone with actions. Both add cost: iterative denoising increases test-time latency, and fine-tuning a large video model is expensive in the low-data regime we target. This raises a natural question: can motion knowledge ina frozen video model be read off directly, without generating future frames?  \nOur starting point is a simple empirical observation: when a frozen text-conditioned Wan 2.2 imageto-video model is queried at the first denoising step, its high-response regions often align with task-relevant contact structure and the robot embodiment. In a bottle manipulation scene, for example, the response highlights the lower graspable region and approaching arm, consistent with both sharper localization and a coarse interaction cue.  \n(c) Benchmark Gains on LIBERO and RoboTwin 2.0  \nFigure 1: KAM-WM provides where-and-how cues for low-data manipulation. (a) KAM highlights interaction-relevant regions and motion cues beyond object masks. (b) These tokens condition the diffusion policy together with RGB and proprioception. (c) KAM-WM improves performance on LIBERO and RoboTwin 2.0 .  \nThis observation is consistent with the Flow Matching training objective [1, 2] . For a conditional video model, the single-step latent velocity at the high-noise endpoint (t=1 .0) estimates a model-dependent displacement toward plausible future video latents cond","cbCaipIHSfAUtgjR","https://ap.wps.com/l/cbCaipIHSfAUtgjR","pdf",6590976,5,1,16,"English","en",105,"# Introduction\n## Motivation: where-and-how priors\n## KAM-WM approach and architecture\n## Experiments on LIBERO and RoboTwin 2.0\n## Contributions and ablations","[{\"question\":\"What problem does KAM-WM address in low-data robot manipulation?\",\"answer\":\"It addresses the need for visual priors that encode both where to interact and how to initiate the interaction, which static cues like segmentation masks do not provide.\"},{\"question\":\"How does KAM-WM derive the Kinematic Affordance Map from a frozen video model?\",\"answer\":\"It queries a Flow Matching image-to-video backbone once and treats the single-step latent velocity at the high-noise endpoint as a Kinematic Affordance Map.\"},{\"question\":\"How are the affordance cues used for control?\",\"answer\":\"A lightweight Perceiver compresses the KAM into tokens that condition a diffusion policy using multi-view RGB observations and proprioception.\"}]",1784190907,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"kam-wm-kinematic-affordance-maps-from-latent-world-models-for-robot-manipulation","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/kam-wm-kinematic-affordance-maps-from-latent-world-models-for-robot-manipulation/83841/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does KAM-WM address in low-data robot manipulation?","Question",{"text":76,"@type":77},"It addresses the need for visual priors that encode both where to interact and how to initiate the interaction, which static cues like segmentation masks do not provide.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does KAM-WM derive the Kinematic Affordance Map from a frozen video model?",{"text":81,"@type":77},"It queries a Flow Matching image-to-video backbone once and treats the single-step latent velocity at the high-noise endpoint as a Kinematic Affordance Map.",{"name":83,"@type":74,"acceptedAnswer":84},"How are the affordance cues used for control?",{"text":85,"@type":77},"A lightweight Perceiver compresses the KAM into tokens that condition a diffusion policy using multi-view RGB observations and proprioception.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":29,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":20,"slug":137},19,"General","general"]