[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84144-en":3,"doc-seo-84144-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84144,2336464648746,"Skyler","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Vision Language Action (VLA) Models for Unmanned Aerial Robotics and Bimanual Manipulation: A Review","Vision–Language–Action (VLA) models integrate visual perception, natural-language understanding, and action generation into a single foundation framework, enabling robots to execute instructions like “fold the towel” or “fly to the red building” from camera observations. With internet-scale pre-training, VLAs support world-knowledge transfer and have become a leading approach for learning-based manipulation. The review surveys 183 contributions (2017–2026) across architectures, training recipes, action representations, bimanual coordination, UAV navigation/control, language grounding, and cross-cutting memory and world-model topics.","Review  \nVision–Language–Action (VLA) Models for Unmanned Aerial Robotics and Bimanual Manipulation: A Review  \nInkyu Sa 1, *, Chanoh Park 2, Hea-Min Lee 3, Donghee Noh 3 and Ho Seok Ahn 4  \nAcademic Editor: Peihu Duan  \nReceived: 9 April 2026  \nRevised: 13 May 2026  \nAccepted: 21 May 2026  \nPublished: 26 May 2026  \nCopyright: © 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.  \narXiv :2607 .06706v 1 [ cs .RO] 7 Jul 2026  \n1 Chef Robotics, San Francisco, CA 94103, USA  \n2 RovifyLab, Gyeonggi 13840, Republic of Korea; [chanoh.park@rovifylab.com](chanoh.park@rovifylab.com)  \n3 IT Application Research Center, Jeonbuk Regional Branch, Korea Electronics Technology Institute (KETI), Jeonju 54853, Republic of Korea; [lee10849@keti.re.kr](lee10849@keti.re.kr) (H.-M.L.); [dhee.noh@keti.re.kr](dhee.noh@keti.re.kr) (D.N.)  \n4 Department of Electrical, Computer and Software Engineering, University of Auckland, Auckland 1010, New Zealand ; [hs.ahn@auckland.ac.nz](hs.ahn@auckland.ac.nz)  \n* Correspondence: [inkyu@chefrobotics.ai](inkyu@chefrobotics.ai)  \nHighlights  \nWhat are the main findings?  \n• The survey finds that VLA research is converging on continuous, chunked action generation—especially flow-matching and hybrid designs—because they avoid the quantization bottleneck of autoregressive action tokens and the latency burden of multistep diffusion, making them better suited to tightly coordinated bimanual control and transferable to aerial systems.  \n• It also finds that progress is driven as much by training strategy as by model architecture: cross-embodiment data diversity and co-training improve downstream generalization more reliably than raw dataset scale alone, while reinforcement learning from autonomous practice is emerging as the key mechanism for surpassing demonstrationlimited performance—that is, achieving higher task success rates, faster execution, and broader generalization than what teleoperated demonstrations alone can support.  \nWhat are the implications of the main findings?  \n• These trends suggest that the most promising path for real-world deployment is not a monolithic end-to-end model but a dual-system design that combines a slower reasoning module with a faster action module, enabling both semantic understanding and highfrequency control in manipulation and aerial robotics.  \n• Looking forward, the field is likely to expand toward production-grade embodied autonomy through end-to-end drone VLAs, aerial manipulation, memory and world-model integration, standardized bimanual benchmarks, safety certification, and continuous self-improvement pipelines that close the gap between benchmark performance and industrial reliability.  \nAbstract  \nVision–Language–Action (VLA) models unify visual perception, natural-language understanding, and action generation within a single foundation model, allowing a robot to follow instructions such as “fold the towel” or “fly to the red building” directly from camera images. Because VLAs inherit world knowledge from internet-scale pre-training, they have become the dominant framework for learning-based manipulation, with bimanual coordination serving as the most demanding testbed: two arms with 7+ degrees of freedom each must move in concert to fold, assemble, and reorient objects. Unmanned aerial robotics faces a structurally similar challenge: a drone must coordinate thrust, attitude, and increasingly gripper commands from visual observations under strict latency and payload constraints. This review covers 183 contributions spanning 2017–2026 and  \norganized along seven dimensions: VLA architectures, training recipes, action representations, bimanual coordination (2022–2026), unmanned aerial vehicle (UAV) navigation and control (2017–2026), language grounding, and cross-cutting concerns including memory and world models. We show that the coor","cbCaiaysKJ05kLFY","https://ap.wps.com/l/cbCaiaysKJ05kLFY","pdf",1479563,5,1,59,"English","en",105,"# Introduction\n## VLA overview and scope\n# Abstract and contribution summary\n## Survey coverage and key dimensions","[{\"question\":\"What challenges make bimanual manipulation an important testbed for VLA models?\",\"answer\":\"Bimanual tasks require two multi-degree-of-freedom arms to coordinate under partial observability, making joint motion, assembly, and reorientation particularly demanding for VLA models.\"},{\"question\":\"How do VLA models generate robot actions from language and images?\",\"answer\":\"A VLA uses a vision–language model to encode visual and language inputs, then produces motor commands through a learned action head, allowing a shared model family to work across different robot embodiments.\"},{\"question\":\"What main research trends does the review identify for improving VLA performance?\",\"answer\":\"The review reports convergence toward continuous, chunked action generation (notably flow-matching and hybrid designs) and emphasizes that training strategy—such as cross-embodiment data diversity and co-training—often improves generalization more reliably than scaling dataset size alone, while reinforcement learning from autonomous practice helps surpass demonstration-limited results.\"}]",1784193372,149,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"vision-language-action-vla-models-for-unmanned-aerial-robotics-and-bimanual-manipulation-a-review","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/vision-language-action-vla-models-for-unmanned-aerial-robotics-and-bimanual-manipulation-a-review/84144/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What challenges make bimanual manipulation an important testbed for VLA models?","Question",{"text":76,"@type":77},"Bimanual tasks require two multi-degree-of-freedom arms to coordinate under partial observability, making joint motion, assembly, and reorientation particularly demanding for VLA models.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How do VLA models generate robot actions from language and images?",{"text":81,"@type":77},"A VLA uses a vision–language model to encode visual and language inputs, then produces motor commands through a learned action head, allowing a shared model family to work across different robot embodiments.",{"name":83,"@type":74,"acceptedAnswer":84},"What main research trends does the review identify for improving VLA performance?",{"text":85,"@type":77},"The review reports convergence toward continuous, chunked action generation (notably flow-matching and hybrid designs) and emphasizes that training strategy—such as cross-embodiment data diversity and co-training—often improves generalization more reliably than scaling dataset size alone, while reinforcement learning from autonomous practice helps surpass demonstration-limited results.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]