[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84142-en":3,"doc-seo-84142-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84142,2336464648746,"Skyler","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Pelican-VLA 0.5 Attending Before Acting Benefits Generalization","Vision-Language-Action (VLA) models aim to generalize across objects, scenes, tasks, and robot embodiments, yet many approaches still rely on task- and environment-specific robot data plus additional data collection and fine-tuning. Pelican-VLA 0.5 analyzes attention behavior and introduces BotTokens to route task-relevant visual information through a learnable bottleneck. The model achieves attention-level generalization in zero-shot settings without annotations or task-specific fine-tuning, supported by persistent, manipulation-centric attention patterns.","arXiv :2607 .06655v2 [ cs .RO] 9 Jul 2026  \nPelican-VLA 0.5:  \nAttending Before Acting Benefits Generalization  \nBeijing Innovation Center of Humanoid Robotics (X-Humanoid)  \nWFM System Group  \n{vito.dai,jian.tang,[jason.ju](jason.ju}@x-humanoid.com)[}](jason.ju}@x-humanoid.com)[@x-humanoid.com](jason.ju}@x-humanoid.com)  \n[https://github. com/Open-X-Humanoid/Pelican-VLA05](https://github. com/Open-X-Humanoid/Pelican-VLA05)  \n[https://huggingface. co/X-Humanoid/Pelican-VLA05](https://huggingface. co/X-Humanoid/Pelican-VLA05)  \nJuly 10, 2026  \nModel  \ninput π0.5 X-VLA GR00T N1 .6 Being-H0 .5 LingBot-VLA ABot-M0 Pelican-VLA  \nReal Sim  \nFranka: Pick up the black bowl between the plate and the ramekin and place it on the table.  \nAgileX Aloha: Use dual arms to pick the drink can and soda can, placing both in the plastic box.  \nAgilex Cobot Magic: Organize the kitchen tableware into the drawer.  \nFigure 1: Attention visualization comparison with open-source VLA baselines in the zero-shot setting. Before taskspecific fine-tuning, Pelican-VLA 0.5 directs its action-pathway attention to the manipulation-relevant object and contact area. In contrast, other open-source VLA models show more diffuse attention, often spreading over the robot arm, surrounding objects, or background.  \nAbstract  \nA central goal of Vision-Language-Action (VLA) research is to build robotic models that can truly generalize across objects, scenes, tasks, and embodiments. However, current VLA models still depend heavily on task-and environment-specific robot data, and typically require additional data collection and fine-tuning when transferred to new objects, scenes, tasks, or embodiments. Recent studies suggest that VLA generalization depends not merely on fitting low-level action trajectories, but also on whether task-relevant information can be represented in a transferable form and effectively routed to the action-generation pathway. This suggests that helping the action pathway attend more consistently to manipulation-relevant regions may be ben-  \neficial for generalization. Motivated by this view, we analyze the attention of VLA models and find that their action pathways attend diffusely, spreading over the robot arm, background, and task-irrelevant objects.  \nIn this report, we present Pelican-VLA 0.5, a unified VLA model that integrates vision-language understanding, futureframe generation, and action prediction within a single architecture. Pelican-VLA 0.5 achieves attention-level generalization: without object annotations, segmentation masks, attention supervision, or task-specific fine-tuning, its action pathway already focuses on the manipulation-relevant object and contact region. This behavior persists across unseen scenes and unseen robot embodiments, and is substantially stronger than in other open-source VLA baselines. We verify that this ability originates from the learnable Bottleneck Tokens (BotTokens) inserted between perception and action: by routing task-relevant visual information through a compact bottleneck, the tokens interface induces manipulation-centric attention during pre-training and remains effective across different policy structures, including a MoT-style architecture.  \nAfter fine-tuning on RoboTwin, Pelican-VLA 0.5 achieves 91.4% success on RoboTwin Clean and 91.0% on RoboTwin Randomized, the best average among open-source VLA baselines. In zero-shot settings, including unseen scenes, unseen objects, and new robot embodiments, Pelican-VLA 0.5 attains non-zero success on several tasks, an early glimmer of generalization. Notably, the attention patterns before and after fine-tuning remain highly similar, suggesting that fine-tuning mainly strengthens the mapping from these pre-formed, manipulation-centric attention regions to executable actions, rather than creating them from scratch.  \nThese findings clarify the meaning of Pelican-VLA 0.5 . The model has already achieved strong attention-level generalization during pre-trai","cbCaiipMenZLyioO","https://ap.wps.com/l/cbCaiipMenZLyioO","pdf",8466488,6,1,17,"English","en",105,"# Abstract\n# Introduction\n# Pelican-VLA 0.5","[{\"question\":\"What problem does Pelican-VLA 0.5 address in VLA generalization?\",\"answer\":\"It targets the gap where VLA models generalize poorly because they often require task/environment-specific robot data and additional fine-tuning when transferred to new objects, scenes, tasks, or embodiments.\"},{\"question\":\"How does Pelican-VLA 0.5 achieve attention-level generalization?\",\"answer\":\"It uses a unified architecture and learnable Bottleneck Tokens (BotTokens) inserted between perception and action, routing task-relevant visual information so the action pathway focuses on manipulation-relevant objects and contact regions during pre-training.\"},{\"question\":\"What evidence suggests the bottleneck drives the learned attention behavior?\",\"answer\":\"Attention patterns already focus on manipulation-relevant regions before fine-tuning and remain highly similar after fine-tuning, indicating fine-tuning primarily strengthens attention-to-action mapping rather than creating the attention regions from scratch.\"}]",1784193359,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"pelican-vla-05-attending-before-acting-benefits-generalization","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/pelican-vla-05-attending-before-acting-benefits-generalization/84142/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does Pelican-VLA 0.5 address in VLA generalization?","Question",{"text":76,"@type":77},"It targets the gap where VLA models generalize poorly because they often require task/environment-specific robot data and additional fine-tuning when transferred to new objects, scenes, tasks, or embodiments.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does Pelican-VLA 0.5 achieve attention-level generalization?",{"text":81,"@type":77},"It uses a unified architecture and learnable Bottleneck Tokens (BotTokens) inserted between perception and action, routing task-relevant visual information so the action pathway focuses on manipulation-relevant objects and contact regions during pre-training.",{"name":83,"@type":74,"acceptedAnswer":84},"What evidence suggests the bottleneck drives the learned attention behavior?",{"text":85,"@type":77},"Attention patterns already focus on manipulation-relevant regions before fine-tuning and remain highly similar after fine-tuning, indicating fine-tuning primarily strengthens attention-to-action mapping rather than creating the attention regions from scratch.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]