[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86231-en":3,"doc-seo-86231-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86231,1374391974585,"Genevieve","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","WALA Learning Executable Latent Actions from Action-Labeled Demonstrations and Action-Free Videos","Generalizable robot policies usually depend on action-labeled demonstrations, but such data are costly and hard to scale, while human videos are abundant yet lack action annotations for direct robot control. WALA presents a framework that jointly learns executable latent actions from both action-labeled demonstrations and action-free videos. It pretrains a semantic-geometric latent action model from videos by learning representations of scene evolution using sparse future observations and predicts future deltas in DINOv3 and dense depth spaces. Policy training freezes the encoder and trains the decoder as a latent world model, with supervision from robot action prediction, latent target matching, and future dynamics prediction. Experiments show strong results on RoboTwin and 75.2% average success on RoboCasa, with additional real-robot validation of generalization across manipulation tasks.","WALA Learning Executable Latent Actions from Action-Labeled Demonstrations and Action-Free Videos  \nJiahao Liu 1,2,3 , Zhongpu Xia2,†, Shuai Tian1,3 , Huangrui Li 1,2,3 , Yuhang Zheng4 , Ning Ma2,5 , Xin Fu2,3 , Xiaotian Liu2 , Jing Li2 , Yixian Li2 , ShangQing Zhou 1,2,3 , Zebin Xing 1,3 , Linbo Wang 1,3 , Chaoyue Li 1,3 ,  \nHaoran Li 1,3,∗ , Dongbin Zhao 1,3,∗  \narXiv :2607 . 11397v1 [ cs .RO] 13 Jul 2026  \nAbstract—Generalizable robot policies typically rely on robot demonstrations with action annotations, yet such data are expensive to collect and difficult to scale. In contrast, largescale and readily available human videos record rich physical interactions, but lack action annotations that can be directly used for robot control. We present WALA, a framework that jointly learns executable latent actions from action-labeled demonstrations and action-free videos. WALA first pretrains a semantic-geometric latent action model on videos without action annotations, enabling it to learn action-relevant representations from scene evolution between the current observation and multiple sparsely sampled future observations. Specifically, WALA forms semantic and geometric future deltas, from which the encoder extracts latent action targets, while the decoder predicts future deltas in the DINOv3 feature space and dense depth space. This avoids raw pixel reconstruction, reducing the influence of appearance details while preserving task-relevant semantic and geometric structure. During policy training, the pretrained encoder remains frozen to providestable latent action targets, while the decoder serves as a trainable latent world model. The latent actions generated by the vision-language backbone are jointly supervised by robot action prediction, latent action target matching, and future dynamics prediction. Action-labeled demonstrations provide both executable control and dynamics supervision, whereas actionfree videos require no robot action labels and still participate in training through latent action and future dynamics supervision. In this way, WALA connects physical scene evolution in videos with executable robot control. Experiments show that WALA achieves strong performance on RoboTwin and reaches an average success rate of 75.2% on RoboCasa, setting a new stateof-the-art result. Additional real-robot experiments further evaluate its generalization ability across diverse manipulation tasks.  \nProject page: WALA Project Page  \nI. INTRODUCTION  \nLearning generalizable robot policies requires data and supervision that cover diverse objects, scenes, task compositions, robot embodiments, and physical interactions. In recent years, vision-language-action models (VLAs) have made significant progress on language-conditioned manipulation by transferring large-scale vision-language pretraining to robot control [7], [8], [9], [10],[11] . Despite this progress, current VLAs still face two coupled limitations. First, they mainly rely on robot demonstrations with action annotations, which  \n∗ Corresponding author.  \n†Project leader.  \n1CASIA, 2Anyverse Dynamics, 3UCAS, 4NUS, 5XJYLU.  \nare expensive to collect, difficult to align across platforms, and insufficient to cover the long tail of physical interactions in the real world. Second, the standard action-supervised training objective primarily learns a mapping from current observations and language instructions to robot actions. It provides only weak and indirect supervision about the future physical consequences of those actions. For long-horizon, contact-rich, or spatially precise manipulation tasks, a policy must not only output plausible motor commands, but also understand how objects, contacts, and scene geometry should evolve after an action is executed.  \nIn contrast, large-scale and readily available human videos record rich physical interactions at much lower cost, including object pose changes, contact events, occlusion relationships, tool use, and goal-directed behavior. These ","cbCaihjaG4aSXd3n","https://ap.wps.com/l/cbCaihjaG4aSXd3n","pdf",5153690,4,1,12,"English","en",105,"# Introduction\n## Problem Motivation\n## Limitations of Existing VLA Training\n## Role of Action-Free Human Videos\n## World Models and Future Dynamics as Training Signals","[{\"question\":\"为什么WALA要同时利用动作标注演示和无动作标注视频？\",\"answer\":\"动作标注演示能提供可执行控制与动力学监督，但收集成本高；无动作视频包含丰富的物理世界演化信息，却缺少可用于低层监督的动作标签。WALA通过把无动作视频也纳入潜在动作与未来动力学监督来补足两者的不足。\"},{\"question\":\"WALA的核心训练目标如何让潜在动作可执行？\",\"answer\":\"WALA用视觉-语言骨干生成潜在动作，并通过机器人动作预测、潜在动作目标匹配以及未来动力学预测对潜在动作进行联合监督，使其同时对控制与未来场景演化产生对齐。\"},{\"question\":\"语义-几何潜在动作模型如何从视频中学习？\",\"answer\":\"WALA在无动作标注视频上预训练语义-几何潜在动作模型：它在当前观测与多个稀疏采样的未来观测之间形成语义与几何的future deltas，然后编码器提取与动作相关的潜在目标，并在DINOv3特征空间与稠密深度空间预测未来deltas，以避免直接像素重建的干扰。\"}]",1784209657,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"wala-learning-executable-latent-actions-from-action-labeled-demonstrations-and-action-free-videos","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/wala-learning-executable-latent-actions-from-action-labeled-demonstrations-and-action-free-videos/86231/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"为什么WALA要同时利用动作标注演示和无动作标注视频？","Question",{"text":75,"@type":76},"动作标注演示能提供可执行控制与动力学监督，但收集成本高；无动作视频包含丰富的物理世界演化信息，却缺少可用于低层监督的动作标签。WALA通过把无动作视频也纳入潜在动作与未来动力学监督来补足两者的不足。","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"WALA的核心训练目标如何让潜在动作可执行？",{"text":80,"@type":76},"WALA用视觉-语言骨干生成潜在动作，并通过机器人动作预测、潜在动作目标匹配以及未来动力学预测对潜在动作进行联合监督，使其同时对控制与未来场景演化产生对齐。",{"name":82,"@type":73,"acceptedAnswer":83},"语义-几何潜在动作模型如何从视频中学习？",{"text":84,"@type":76},"WALA在无动作标注视频上预训练语义-几何潜在动作模型：它在当前观测与多个稀疏采样的未来观测之间形成语义与几何的future deltas，然后编码器提取与动作相关的潜在目标，并在DINOv3特征空间与稠密深度空间预测未来deltas，以避免直接像素重建的干扰。","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]