[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82242-en":3,"doc-seo-82242-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82242,962075114765,"Quinn","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Causally Debiased Latent Action Model for Embodied Action Conditioned World Models","Action-conditioned world models (ACWMs) generate future observations conditioned on embodied actions, supporting robot planning, policy evaluation, and data augmentation. Real-world action-labeled data is expensive, motivating latent action models (LAMs) that infer action representations from videos without executable labels. Reconstruction-only LAM training entangles action-relevant dynamics with action-irrelevant visual factors, creating controllability bias and fragility. CD-LAM introduces three finetuning debiasing objectives—embodiment-centric reconstruction, action-centric contrastive learning, and latent calibration—to improve latent-action controllability, action following, visual fidelity, and adaptation efficiency on DreamDojo backbones.","Causally Debiased Latent Action Model for Embodied Action Conditioned World Models  \nYufan Wei 1,2∗, Kun Zhou 1†, Lingjun Mao 1,2∗, Zijun Zhang 1 , Ziming Xu 1 , Ziqiao Xi 1 , Shuang Liang 1,2∗, Ruobing Han 1 , Yuchen Yan 1 , Xinyue Wang 1,2∗, Fan Feng 1 , Biwei Huang 1  \n1Aether AI 2University of California, San Diego  \n∗Done during an internship at Aether AI.  \n†Project Lead & Corresponding author: [franciskunzhou@gmail.com](franciskunzhou@gmail.com)  \narXiv :2607 .09185v1 [ cs .CV] 10 Jul 2026  \n2B  \n14B  \n2B  \n14B  \n2B  \n14B  \nFine-tuning  \nAdaptation  \n(a) Action Following (b) Visual Fidelity (c) Efficiency (d) Causal Analysis  \nFig. 1. Performance overview and the underlying confounding mechanism. CD-LAM substantially improves action following, visual fidelity, and data efficiency on DreamDojo. (a) CD-LAM lowers embodiment action-following error (FDCE) at both 2B and 14B. (b) CD-LAM also raises PSNR at both scales (gains annotated in dB) . (c) CD-LAM uses 3k and 6k robot action adaptation updates for the final 2B and 14B checkpoints, compared with DreamDojo’s 50k-update reference; crossing curves in Fig. 8. (d) Causal analysis: a reconstruction-trained LAM can leak action-irrelevant confounders into the action condition; CD-LAM reduces the measured shortcut dependence along this path, and (a)–(c) quantify the resulting gains.  \nAbstract—Action-conditioned world models (ACWMs) aim to simulate future observations conditioned on embodied actions, offering a promising foundation for robot planning, policy evaluation, and data augmentation. However, learning controllable ACWMs requires large-scale action-labeled data, which remains costly to collect in the real world. Latent action models (LAMs) mitigate this bottleneck by inferring latent actions from videos without executable action labels, but existing LAMs are typically trained with reconstruction-only objectives and therefore entangle action-relevant dynamics with action-irrelevant visual factors such as backgrounds and non-interacted objects. In this work, we identify this action-irrelevant bias as a key obstacle to controllable ACWMs and introduce evaluation metrics to measure latentaction bias, action following, and robustness. We propose CDLAM, a causally debiased framework for LAM-based ACWMs. CD-LAM introduces three debiasing objectives used during finetuning: embodiment-centric reconstruction, action-centric contrastive learning, and latent space calibration, which together encourage embodiment-focused, action-aware, and well-calibrated, non-collapsed latent action representations. Experiments on 2Band 14B ACWM backbones show that CD-LAM substantially improves latent-action controllability, downstream robot action following, visual fidelity, and adaptation efficiency: at 14B, CDLAM matches the DreamDojo reference with more than 12× fewer robot action adaptation updates and surpasses it at the 6k  \nfinal checkpoint.  \nIndex Terms—latent action models, world models, causal debiasing  \nI. INTRODUCTION  \nAction conditioned world models (ACWMs) [1]–[4] have emerged as a promising paradigm for simulating the physical world directly from visual observations and embodied control signals. Given the current visual observation and an action trajectory, an ACWM aims to forecast the resulting future observation sequence, thereby simulating how the environment and embodiment would evolve under that intervention. By predicting the futures of different candidate action trajectories, ACWMs have the potential to serve as general-purpose simulators for robot planning, policy evaluation, and data augmentation.  \nHowever, the controllability of ACWMs relies on largescale action-labeled data, whereas collecting robot videos with action annotations remains costly in the real world [5], [6] . In  \nForeground Background  \nLatent Action Decoder  \n􀀡' (􀀣 + 1)  \nLatent Action Model  \nLAM Debiased Fine-tuning  \nFig. 2. Overview of CD-LAM. CD-LAM debiases the LAM’s latent action space in three ","cbCaidFQsY7HWaQU","https://ap.wps.com/l/cbCaidFQsY7HWaQU","pdf",9910486,2,1,14,"English","en",105,"# Introduction\n# CD-LAM Overview\n## Debiasing stages and objectives\n# Problem: Action-irrelevant confounding\n## Reconstruction-only bias and its effects","[{\"question\":\"Why are action-conditioned world models difficult to train in practice?\",\"answer\":\"They depend on large-scale action-labeled data, while collecting robot videos with actionable action annotations in the real world is costly.\"},{\"question\":\"What limitation of existing latent action models causes biased controllability?\",\"answer\":\"Training with reconstruction-only objectives allows action-irrelevant visual factors (e.g., backgrounds and non-interacted objects) to leak into the latent action representation, confounding action conditioning.\"},{\"question\":\"How does CD-LAM reduce action-irrelevant bias during fine-tuning?\",\"answer\":\"CD-LAM applies three debiasing objectives: embodiment-centric reconstruction, action-centric contrastive learning, and latent space calibration, which encourage embodiment-focused, action-aware, and well-calibrated latent action representations.\"}]",1784179080,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"causally-debiased-latent-action-model-for-embodied-action-conditioned-world-models","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/causally-debiased-latent-action-model-for-embodied-action-conditioned-world-models/82242/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-19","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why are action-conditioned world models difficult to train in practice?","Question",{"text":75,"@type":76},"They depend on large-scale action-labeled data, while collecting robot videos with actionable action annotations in the real world is costly.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What limitation of existing latent action models causes biased controllability?",{"text":80,"@type":76},"Training with reconstruction-only objectives allows action-irrelevant visual factors (e.g., backgrounds and non-interacted objects) to leak into the latent action representation, confounding action conditioning.",{"name":82,"@type":73,"acceptedAnswer":83},"How does CD-LAM reduce action-irrelevant bias during fine-tuning?",{"text":84,"@type":76},"CD-LAM applies three debiasing objectives: embodiment-centric reconstruction, action-centric contrastive learning, and latent space calibration, which encourage embodiment-focused, action-aware, and well-calibrated latent action representations.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]