[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84778-en":3,"doc-seo-84778-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84778,5909877438554,"Maeve","https://ap-avatar.wpscdn.com/avatar/5600025385ad2bf12a7?_k=1778553567797529272",8,"Research & Report","InternVLA-A1.5 Unifying Understanding Latent Foresight and Action for Compositional Generalization","InternVLA-A1.5 presents a unified robotics policy that combines semantic understanding from pretrained vision-language models with physical dynamics learned through future prediction, addressing common failures in existing unified VLA designs. The method keeps a native VLM backbone trained on VQA and subtask prediction, adds a lightweight expert for continuous action generation, and reformulates future prediction as latent querying via learnable foresight tokens supervised by a frozen pretrained video generation model. A video branch is removed at inference for real-time control. Trained on 1.2M robot episodes and 3M multimodal samples, it reaches top performance on six simulation benchmarks and shows stronger compositional generalization in real-world tests.","arXiv :2607 .04988v 1 [ cs .RO] 6 Jul 2026  \nInternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization  \nPhysical Intelligence Team, Shanghai AI Laboratory  \nFull author list in Contributors section  \nUnified models for robot manipulation aim to equip one policy with both the semantic priors of pretrained VLMs and the physical dynamics learned through future prediction. In practice, existing designs tend to erode the semantics of thepretrained backbone, suffer interference among heterogeneous objectives, and learn future prediction from scratch in pixel space, leaving the dynamics priors of pretrained video generators unexploited. We present InternVLA-A1.5, which builds the policy on a native VLM backbone that keeps training on VQA and subtask prediction, and attaches a lightweight unified expert for continuous action generation. Future prediction is recast as a latent-querying problem, where a small set of learnable foresight tokens condenses the task-relevant future into a compact latent code under the supervision of afrozen pretrained video generation model, so the policy inherits world-model dynamics priors without ever learning pixel-level generation. The video branch is discarded at inference, keeping real-time control. Pretrained on 1. 2M robot episodes and 3M multimodal samples, InternVLA-A1. 5 achieves the best overall resultson all six simulation benchmarks. In the real world, the preserved semantics deliver the strongest compositional generalization on held-out instruction bindings, and the two designs together sustain long-horizon execution.  \nÑ Homepage | § Code:InternVLA-A1 .5 | Æ Model:InternVLA-A1 .5  \nInternVLA-A1.5 Unifying Understanding, Latent Foresight, and Action for Compositional Generalization  \nSimulation Benchmark Results  \n93.2  \n98.9 35.2  \nInstruction Following  \n 􀀡 !. \\#  Motus  InternVLA-A1 .5  \nSort Tubes Insert Tubes Move Tubes  \nFigure 1 . Overview of InternVLA-A1.5. InternVLA-A1.5 unifies understanding, latent foresight, and action by attaching a lightweight expert to a pretrained VLM backbone. It is co-trained on vision-language and robot manipulation data, and introduces learnable foresight tokens to extract task-relevant future information as a compact latent representation. This representation is supervised by a frozen video generation model, which is used only during training and removed at inference. Extensive simulation and real-world experiments validate the effectiveness of the design.  \n1. Introduction  \nDeveloping general-purpose robots that can manipulate diverse objects under language instructions is a long-standing goal of embodied intelligence, with broad application value in homes, factories, and other unstructured environments. The strong generalization of vision-language models (Achiamet al., 2023; Beyer et al., 2024; Qwen Team, 2026) and the powerful generative ability of video generation models (Agarwal et al., 2025; Wan et al., 2025) have motivated researchers to explore transferring such capabilities to embodied control. This has given rise to Vision-Language-Action (VLA) models (Bjorck et al., 2025; Intelligence et al., 2024, 2025a,b, 2026; Yu et al., 2026), which inherit rich semantic priors from pretrained vision-language backbones and generalize well across objects and instructions. In parallel, Video Action and World Action models (Bi et al., 2025; GEAR, 2026; Kim et al., 2026; Li et al., 2026; Team et al., 2026; Zhou et al., 2026) learn to predict future visual states and thus acquire a strong sense of physical dynamics. This complementarity has motivated a growing body of work to explore whether a single unified model (Cai et al., 2026a; Hu et al., 2026; Liu et al., 2026; Lu et al., 2025; Luo et al., 2026; Sun et al., 2026) can combine the strengths of both. Along this line, our prior work InternVLA-A1 (Cai et al., 2026a) makes an early attempt by treating future visual states and actions jointly as training targets with","cbCaimuL5SFUvydb","https://ap.wps.com/l/cbCaimuL5SFUvydb","pdf",6903653,1,24,"English","en",105,"# Introduction\n## Motivation and background\n## Limitations of existing unified models\n## Key idea of InternVLA-A1.5","[{\"question\":\"What problem does InternVLA-A1.5 address in compositional robotic generalization?\",\"answer\":\"It targets how unified VLA models can degrade semantic understanding, suffer interference among heterogeneous objectives, and fail to exploit pretrained dynamics without resorting to expensive pixel-level video generation.\"},{\"question\":\"How does InternVLA-A1.5 preserve semantic abilities during training?\",\"answer\":\"It keeps training a native vision-language backbone with VQA and subtask prediction, and adds a discrete action token objective to provide action-aware supervision while strengthening instruction following.\"},{\"question\":\"What is the role of latent foresight tokens and the frozen video generation model?\",\"answer\":\"Future prediction is converted into latent querying: learnable foresight tokens condense task-relevant future into a compact latent code supervised by a frozen pretrained video generator. The video branch is used only during training and removed at inference for real-time control.\"}]",1784198172,60,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"internvla-a15-unifying-understanding-latent-foresight-and-action-for-compositional-generalization","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/internvla-a15-unifying-understanding-latent-foresight-and-action-for-compositional-generalization/84778/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does InternVLA-A1.5 address in compositional robotic generalization?","Question",{"text":75,"@type":76},"It targets how unified VLA models can degrade semantic understanding, suffer interference among heterogeneous objectives, and fail to exploit pretrained dynamics without resorting to expensive pixel-level video generation.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does InternVLA-A1.5 preserve semantic abilities during training?",{"text":80,"@type":76},"It keeps training a native vision-language backbone with VQA and subtask prediction, and adds a discrete action token objective to provide action-aware supervision while strengthening instruction following.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the role of latent foresight tokens and the frozen video generation model?",{"text":84,"@type":76},"Future prediction is converted into latent querying: learnable foresight tokens condense task-relevant future into a compact latent code supervised by a frozen pretrained video generator. The video branch is used only during training and removed at inference for real-time control.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":28,"slug":108},5,"Comic","comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]