[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85106-en":3,"doc-seo-85106-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85106,687197207057,"Sage","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Native Video Action Pretraining for Generalizable Robot Control","Video-action models promise strong robot control, yet reusing video-generation architectures designed for digital content can fail in physical environments. LingBot-VA 2.0 introduces a native video-action foundation model built for embodiment. It uses a semantic visual-action tokenizer, causal pretraining from scratch, a sparse MoE backbone for efficient high-rate inference, and an enhanced asynchronous closed-loop scheme with learned forward dynamics. Real-world deployment shows robust few-shot generalization across complex manipulation tasks.","arXiv :2607 .08639v 1 [ cs .RO] 9 Jul 2026  \nNative Video-Action Pretraining for Generalizable Robot Control  \nQihang Zhang, Lin Li, Luyao Zhang, Shuai Yang, Yiming Luo, Shuaiting Li, Ruilin Wang, Junke Wang, Jiahao Shao, Gangwei Xu, Jiaming Zhou, Yishu Shen, Yudong Jin, Fangyi Xu, Shuailei Ma, Jiaqi Liao, Guanxing Lu, Zifan Shi, Yongkun Wen, Yujie Zhao, Weixuan Tang, Xinyang Wang, Chaojian Li,  \nJiapeng Zhu, Ka Leong Cheng, Nan Xue, Xing Zhu, Yujun Shen, Yinghao Xu†  \n†Project Lead  \nThe advent of video-action models offers a promising path for robot control. Nevertheless, we argue that repurposing video generative models designed for digital content creation is inherently inadequate for physical environments. To bridge this gap, we present LingBot-VA 2.0, a video-action foundation model built from the ground up for embodiment. Four core design principles showcase its evolution from LingBot-VA. (1) Departing from traditional reconstruction-focused VAEs, we introduce a semantic visual-action tokenizer, which aligns visual representations with both semantics and actions, improving instruction following and action precision in subsequent policy learning. (2) Given the strictly causal nature of temporal dynamics, we adopt a causal pretraining paradigm, training from scratch to circumvent the catastrophic forgetting that frequently occurs when adapting bidirectional architectures. (3) To meet the demands of high-frequency inference, our model employs a sparse MoE backbone, expanding model capacity without compromising efficiency. (4) Real-time closed-loop control is realized through an enhanced asynchronous inference scheme, which predicts future latents in parallel with action execution while re-grounding each rollout on the latest observation via learned forward dynamics. Real-world deployment validates LingBot-VA 2.0 as a robust foundation model, as evidenced by its few-shot generalization across complex manipulation tasks.  \nWebsite: [https://technology.robbyant.com/lingbot-va-v2](https://technology.robbyant.com/lingbot-va-v2)  \n1 Introduction  \nVideo-action models such as LingBot-VA [50] and DreamZero [115] have recently emerged as a powerful paradigm for generalist robot manipulation. Rather than mapping observations directly to actions, as reactive vision-language-action policies do [9, 11, 43], they jointly predict how a scene will evolve and how to act within it, grounding control in physical dynamics and improving sample efficiency and generalization [35, 54, 132] . Much of this capability, however, is inherited from their video-pretrained backbones, suggesting that generalist robot control depends as much on thepretrained foundation as on the policy learned upon it.  \nCurrent video-action models are still largely built from components designed for generic video generation—areconstruction-oriented VAE and a bidirectional video-diffusion backbone—with an action module added for robotics afterward. This starting point creates three concrete limitations. First, the representation is optimized for appearance rather than dynamics: pixel-reconstruction latents preserve visual detail but carry limited semantic and physical structure, and the separately attached action module leaves world states and actions in poorly aligned spaces. Second, inference is too slow for closed-loop control: high-dimensional video tokens and iterative denoising make first-generation videoaction models costly to run at the frequencies real robots require. Third, the pretraining signal does not scale toward control: web-scale video is abundant, but generic video objectives do not teach how actions reshape the world, so the action signal remains tied to expensive robot data, limiting the control knowledge learned before downstream adaptation.  \nThese limitations are compounded by a structural mismatch: the backbone is pretrained with bidirectional attention,  \nwhile closed-loop control unfolds strictly forward in time. LingBot-VA resolves this mismatch ","cbCaio5H2Z1cBuwH","https://ap.wps.com/l/cbCaio5H2Z1cBuwH","pdf",10997354,1,29,"English","en",105,"# Introduction\n## Key limitations of existing video-action models\n## Native pretraining strategy and motivation\n## LingBot-VA 2.0 design overview","[{\"question\":\"Why is repurposing digital video generative models inadequate for robot control?\",\"answer\":\"Digital video generators optimize for appearance-focused reconstruction and bidirectional temporal modeling, which misaligns dynamics with actions needed for physical environments. Their inference and control-oriented learning signals also do not match real-time closed-loop requirements.\"},{\"question\":\"What core components does LingBot-VA 2.0 introduce?\",\"answer\":\"LingBot-VA 2.0 combines a semantic visual-action tokenizer that aligns latents with semantics and actions, and a causal diffusion transformer that predicts future visual latents and connects them with latent actions. It conditions on language instructions via cross-attention to a pretrained language model.\"},{\"question\":\"How does LingBot-VA 2.0 enable efficient closed-loop control in real time?\",\"answer\":\"It adopts a causal pretraining paradigm to match strictly forward temporal dynamics, uses a sparse MoE backbone to expand capacity without sacrificing efficiency, and runs an enhanced asynchronous inference scheme that predicts future latents while re-grounding each rollout on the latest observation using learned forward dynamics.\"}]",1784201136,73,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"native-video-action-pretraining-for-generalizable-robot-control","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/native-video-action-pretraining-for-generalizable-robot-control/85106/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is repurposing digital video generative models inadequate for robot control?","Question",{"text":75,"@type":76},"Digital video generators optimize for appearance-focused reconstruction and bidirectional temporal modeling, which misaligns dynamics with actions needed for physical environments. Their inference and control-oriented learning signals also do not match real-time closed-loop requirements.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What core components does LingBot-VA 2.0 introduce?",{"text":80,"@type":76},"LingBot-VA 2.0 combines a semantic visual-action tokenizer that aligns latents with semantics and actions, and a causal diffusion transformer that predicts future visual latents and connects them with latent actions. It conditions on language instructions via cross-attention to a pretrained language model.",{"name":82,"@type":73,"acceptedAnswer":83},"How does LingBot-VA 2.0 enable efficient closed-loop control in real time?",{"text":84,"@type":76},"It adopts a causal pretraining paradigm to match strictly forward temporal dynamics, uses a sparse MoE backbone to expand capacity without sacrificing efficiency, and runs an enhanced asynchronous inference scheme that predicts future latents while re-grounding each rollout on the latest observation using learned forward dynamics.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]