[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82029-en":3,"doc-seo-82029-105":31,"detail-sidebar-cat-0-en-105":93},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},82029,7971461740886,"Theodore","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","FabriVLA A Lightweight Vision Language Action Model for Precise Multi-Task Manipulation","FabriVLA is a lightweight Vision-Language-Action (VLA) model targeting precise multi-task robotic manipulation. The method integrates an InternVL3.5 vision-language backbone with a flow-matching action head that applies gated self-attention across action tokens and uses shallow VLM layer fusion to enhance spatial context. Training follows single-stage joint optimization from a pretrained VLM with a randomly initialized action head. On the Meta-World MT50 benchmark, FabriVLA reaches 90.0% tier-average success, achieving strong results with a 1B-scale backbone.","FabriVLA: A Lightweight Vision-Language-Action Model for Precise Multi-Task  \nManipulation  \nShiyuan Yang 1 , ∗ , Borong Zhang 1 ,∗ , Jizheng Zhang 1 ,∗ , Zhijia Tao 1 ,∗ , Junfei Guo2 , Donglai Ran3 , Xu Bian3 ,†, Qingbiao Li 1 ,†  \n1University of Macau  \n2Mese Technology Limited Co., Ltd., 3FabriX team at Youibot Robotics Co., Ltd.  \narXiv :2607 .08575v2 [ cs .RO] 10 Jul 2026  \nAbstract  \nWe present FabriVLA, a lightweight Vision-Language-Action model for Precise Multi-Task Manipulation. FabriVLA combines an InternVL3.5 vision-language backbone with a flowmatching action head featuring gated self-attention across action tokens and shallow VLM layer fusion for enriched spatial context. The model is trained via single stage joint optimization from a pretrained VLM and randomly initialized action head. On the Meta-World MT50 benchmark spanning 50 diverse manipulation tasks, FabriVLA achieves a tier-average success rate of 90.0%, demonstrating that a compact VLA built on a 1B scale VLM can achieve strong performance without relying on multi billion parameter VLA backbones.  \n1 Introduction  \nVision-Language-Action (VLA) models have emerged as a promising paradigm for generalist robot manipulation, leveraging pretrained vision-language models to ground language instructions in visual observations and produce executable action sequences (Zitkovich et al. 2023; Kim et al. 2024; Black et al. 2026) . While large-scale VLAs with tens of billions of parameters achieve impressive results, their computational cost and inference latency pose practical challenges for real-time robotic control. This motivates the development of lightweight VLA architectures that balance performance with efficiency.  \nIn this work, we present FabriVLA, a lightweight VLA model that achieves competitive performance on the MetaWorld MT50 benchmark (Yu et al. 2020) while maintaining a modest parameter footprint. FabriVLA is inspired by Evo- 1 (Lin et al. 2026c) and builds on lightweight VLA design with an InternVL3.5 (Wang et al. 2025) vision-language backbone, a gated self-attention mechanism within the flowmatching action head, and shallow VLM layer fusion for improved spatial context. Representative camera observations are shown in Figure 1 .  \nThe gated self-attention allows action tokens within the prediction horizon to attend to each other through a learnable gating parameter initialized to zero. At initialization, the gate is closed and the block behaves like a transformer with cross attention only. During training, the gate gradually opens, allowing the model to learn inter-step dependencies among action tokens. This design provides a smooth optimization  \n∗These authors contributed equally.  \n†Corresponding Author.  \nCopyright © 2027, Association for the Advancement of Artificial Intelligence ([www.aaai.org](www.aaai.org)). All rights reserved.  \nFigure 1: Representative Meta-World visual observations from the available camera views. The final FabriVLA setting uses the corner RGB view as its policy input.  \npath while enabling the model to capture the temporal structure of manipulation trajectories. In parallel, shallow VLM layer fusion combines the final VLM layer with an intermediate layer to expose both semantic context and lower-level spatial detail to the action head, which is important for precise object localization and contact rich manipulation.  \nFabriVLA is trained in a single stage with the entire model jointly optimized from the start. On Meta-World MT50, FabriVLA achieves a tier-average success rate of 90.0% and an overall episode-level success rate of 92.0%, with strong performance across all four difficulty tiers: easy (95.0%), medium (88.2%), hard (86.7%), and very hard (90.0%) .  \n2 Model Architecture  \nFabriVLA consists of three components: (1) an InternVL3.5 vision-language backbone,(2) state and action encoders, and (3) a flow-matching action head. Figure 2 illustrates the overall architecture.  \n2.1 Vision-Language Backbone  \nWe ado","cbCaipk4in6KBvL8","https://ap.wps.com/l/cbCaipk4in6KBvL8","pdf",1203681,9,1,7,"English","en",105,"# Introduction\n# Model Architecture\n## Vision-Language Backbone\n## State and Action Encoding\n## Action-Head Modules","[{\"question\":\"What problem does FabriVLA address in robotic manipulation?\",\"answer\":\"FabriVLA targets the need for real-time feasible VLA models by providing strong performance with a lightweight architecture instead of relying on very large VLA backbones.\"},{\"question\":\"How is the model built at a high level?\",\"answer\":\"FabriVLA combines an InternVL3.5 vision-language backbone, state and action encoders, and a flow-matching action head with gated cross-attention blocks.\"},{\"question\":\"What results does FabriVLA achieve on the Meta-World MT50 benchmark?\",\"answer\":\"FabriVLA achieves 90.0% tier-average success and 92.0% overall episode-level success, with reported performance across easy, medium, hard, and very hard tiers.\"}]","FabriVLA A Lightweight Vision Language Action Model for Precise Multi-Task Manipulation | PDF",1784177684,18,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":88,"head_meta":90,"extra_data":92,"updated_unix":29},"fabrivla-a-lightweight-vision-language-action-model-for-precise-multi-task-manipulation","",{"@graph":37,"@context":87},[38,55,70],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":54},"https://docshare.wps.com/document/fabrivla-a-lightweight-vision-language-action-model-for-precise-multi-task-manipulation/82029/",4,{"url":53,"name":13,"@type":56,"author":57,"headline":13,"publisher":59,"fileFormat":62,"inLanguage":24,"description":14,"dateModified":63,"datePublished":64,"encodingFormat":62,"isAccessibleForFree":65,"interactionStatistic":66},"DigitalDocument",{"name":9,"@type":58},"Person",{"url":42,"name":60,"@type":61},"DocShare","Organization","application/pdf","2026-07-29","2026-07-16",true,{"@type":67,"interactionType":68,"userInteractionCount":20},"InteractionCounter",{"@type":69},"ViewAction",{"@type":71,"mainEntity":72},"FAQPage",[73,79,83],{"name":74,"@type":75,"acceptedAnswer":76},"What problem does FabriVLA address in robotic manipulation?","Question",{"text":77,"@type":78},"FabriVLA targets the need for real-time feasible VLA models by providing strong performance with a lightweight architecture instead of relying on very large VLA backbones.","Answer",{"name":80,"@type":75,"acceptedAnswer":81},"How is the model built at a high level?",{"text":82,"@type":78},"FabriVLA combines an InternVL3.5 vision-language backbone, state and action encoders, and a flow-matching action head with gated cross-attention blocks.",{"name":84,"@type":75,"acceptedAnswer":85},"What results does FabriVLA achieve on the Meta-World MT50 benchmark?",{"text":86,"@type":78},"FabriVLA achieves 90.0% tier-average success and 92.0% overall episode-level success, with reported performance across easy, medium, hard, and very hard tiers.","https://schema.org",{"og:url":53,"og:type":89,"og:title":13,"og:site_name":60,"og:description":14},"article",{"robots":91,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":94},[95,99,103,107,112,117,121,124,128,131,135],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":96,"show_sort_weight":97,"slug":98},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},"Literature",80,"literature",{"id":54,"doc_module":4,"doc_module_name":47,"category_name":104,"show_sort_weight":105,"slug":106},"Exam",70,"exam",{"id":108,"doc_module":4,"doc_module_name":47,"category_name":109,"show_sort_weight":110,"slug":111},5,"Comic",60,"comic",{"id":113,"doc_module":4,"doc_module_name":47,"category_name":114,"show_sort_weight":115,"slug":116},6,"Technology",50,"technology",{"id":22,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":125,"show_sort_weight":126,"slug":127},"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":47,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":47,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":47,"category_name":137,"show_sort_weight":108,"slug":138},19,"General","general"]