[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85594-en":3,"doc-seo-85594-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85594,3848291630094,"Emma Wilson","https://eur-avatar.wpscdn.com/davatar_085a072bc5b1113ac321206ff7593b45",8,"Research & Report","IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation","Robot imitation data are often multimodal: visually similar language-conditioned observations can lead to different action chunks because human demonstrators follow different short-horizon intents, task phases, or recent context. Existing frame-conditioned VLA policies infer each chunk from only the current observation and instruction, which under partial observability may resample competing intents across adjacent replanning steps, causing conflicts and unstable execution. IntentVLA is a history-conditioned VLA framework that encodes recent visual observations into a compact short-horizon intent representation to condition chunk generation. An ambiguity-aware benchmark, AliasBench, evaluates this failure mode on RoboTwin2 with matched training data and isolated aliasing environments, improving stability and surpassing strong baselines.","IntentVLA: Short-Horizon Intent Modeling for Aliased Robot  \nManipulation  \nShijie Lian1,2* Bin Yu2,4* Xiaopeng Lin5,2* Zhaolong Shen2,6* Laurence Tianruo Yang1,7,† Yurun Jin3,9 Haishan Liu2 Changti Wu2,8 Hang Yuan2,8 Cong Huang2,3 Kai Chen2,3,10,†  \n1HUST 2ZGCA 3ZGCI  \n4HIT 5HKUST(GZ) 6BUAA 7ZZU 8ECNU 9USTC 10DeepCybo  \narXiv :2605 . 14712v2 [ cs .RO] 11 Jul 2026  \nAbstract  \nRobot imitation data are often multimodal: similar visual-language observations may be followed by different action chunks because human demonstrators act with different shorthorizon intents, task phases, or recent context. Existing frame-conditioned VLA policies infer each chunk from the current observation and instruction alone, so under partial observability they may resample different intents across adjacent replanning steps, leading to inter-chunk conflict and unstable execution. We introduce IntentVLA, a history-conditioned VLA framework that encodes recent visual observations into a compact short-horizon intent representation and uses it to condition chunk generation. We further introduce AliasBench, a 12-task ambiguity-aware benchmark on RoboTwin2 with matched training data and evaluation environments that isolate short-horizon observation aliasing. Across AliasBench, SimplerEnv,  \nLIBERO, and RoboCasa, IntentVLA improves rollout stability and outperforms strong VLA baselines. The benchmark code, generated data, and model code are released at [https:](https:)//[github.com/ZGC-EmbodyAI/IntentVLA](github.com/ZGC-EmbodyAI/IntentVLA).  \n1 Introduction  \nVision-language-action (VLA) models provide a direct interface from perception and instruction to control: given visual observations and a language command, the policy outputs robot actions (Kim et al., 2024 ; Black et al., 2024 ; Liu et al., 2025 ; Bjorck et al., 2025) . Recent large-model-based VLAs scale this paradigm with transformer backbones, large robot datasets, and vision-language pretraining, enabling more generalist manipulation policies across tasks and embodiments (In-  \n*Equal contribution  \n†Corresponding authors  \nWork done at Zhongguancun Academy (Beijing) .  \ntelligence et al., 2025 ; GEAR-Team et al., 2025 ; Bi et al., 2025 ; Zheng et al., 2025a ; Liu et al., 2026) .  \nTraining VLA models typically relies on largescale human-collected robot trajectories (O’Neillet al., 2024 ; Bu et al., 2025a ; Walke et al., 2023 ; Khazatsky et al., 2024), and these datasets often faithfully reflect the underlying multimodality of manipulation behavior. For instance, an environment may admit multiple valid goals, and even a fixed goal can often be achieved through multiple feasible paths (Zhai et al., 2025a) . This diversity is not itself the problem. Human demonstrations are naturally multimodal across episodes, but they are locally committed within each episode: once a demonstrator follows a particular task phase, path, or completion strategy, adjacent action chunks usually remain consistent with that choice. The difficulty arises because current VLA policies generally infer actions from only the current frame image and the language instruction. Under partial observability, the same frame-level observation can correspond to different short-horizon intents, but a frame-conditioned VLA does not observe the episode-level commitment that selected one of them. Repeated chunk generation can then switch among intents across adjacent decision steps, producing contradictory chunks and unstable execution. Thus, the goal is not to eliminate multimodality, but to condition generation on the commitment already expressed by the current episode.  \nFigure 1 illustrates this ambiguity: similar breadholding observations can require skillet placementor plate return under the same instruction. To measure this failure mode, we build AliasBench on RoboTwin2 (Chen et al., 2025) with matched simulation training data and evaluation environments that isolate short-horizon observation aliasing. Rather than only scoring task","cbCaiaRMaNfr0w3k","https://ap.wps.com/l/cbCaiaRMaNfr0w3k","pdf",1816829,2,1,17,"English","en",105,"# Introduction\n## Failure mode: frame-conditioned intent ambiguity under partial observability\n## AliasBench benchmark construction\n## IntentVLA approach: history-conditioned short-horizon intent representation\n## Experimental evaluation and contributions","[{\"question\":\"Why do frame-conditioned VLA chunk policies become unstable under partial observability?\",\"answer\":\"When the same frame-level observation can correspond to different short-horizon intents, a frame-conditioned policy may switch among intents across adjacent replanning steps without observing the episode-level commitment, leading to contradictory chunks and unstable execution.\"},{\"question\":\"What is AliasBench and what problem does it measure?\",\"answer\":\"AliasBench is a 12-task ambiguity-aware benchmark on RoboTwin2 designed to evaluate VLA behavior under short-horizon observation aliasing, testing whether policies preserve local continuity when visually similar states require different next action chunks.\"},{\"question\":\"How does IntentVLA address short-horizon intent aliasing?\",\"answer\":\"IntentVLA conditions chunk generation on recent visual evidence by encoding recent observations into a compact short-horizon intent representation and fusing it with current vision-language context via gated cross-attention, then using this fused context to condition the action head.\"}]",1784204811,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"intentvla-short-horizon-intent-modeling-for-aliased-robot-manipulation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/intentvla-short-horizon-intent-modeling-for-aliased-robot-manipulation/85594/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do frame-conditioned VLA chunk policies become unstable under partial observability?","Question",{"text":75,"@type":76},"When the same frame-level observation can correspond to different short-horizon intents, a frame-conditioned policy may switch among intents across adjacent replanning steps without observing the episode-level commitment, leading to contradictory chunks and unstable execution.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is AliasBench and what problem does it measure?",{"text":80,"@type":76},"AliasBench is a 12-task ambiguity-aware benchmark on RoboTwin2 designed to evaluate VLA behavior under short-horizon observation aliasing, testing whether policies preserve local continuity when visually similar states require different next action chunks.",{"name":82,"@type":73,"acceptedAnswer":83},"How does IntentVLA address short-horizon intent aliasing?",{"text":84,"@type":76},"IntentVLA conditions chunk generation on recent visual evidence by encoding recent observations into a compact short-horizon intent representation and fusing it with current vision-language context via gated cross-attention, then using this fused context to condition the action head.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]