[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83236-en":3,"doc-seo-83236-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},83236,962075114101,"Seraphina","https://ap-avatar.wpscdn.com/avatar/e000253a75eb197efd?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780044092746381165",8,"Research & Report","Multi-Agent Robotic Control with Onboard Vision-Language Models","Vision Language Models (VLMs) and Vision Language Action (VLA) can support robotic control, but they often struggle with explainability, generalization beyond out-of-distribution settings, and heavy compute requirements, commonly tied to network connectivity. The paper proposes a Multi-Agent System (MAS) architecture that runs fully on onboard hardware, using compact VLMs (3–20B) and fine-tuning for package inspection accuracy. A “Megamind” orchestration agent addresses context retention limits in long-horizon planning with smaller models. Validation is performed in a hardware-in-the-loop simulated industrial warehouse, and the simulation environment is released as open source under Apache 2.0.","arXiv :2607 .07403v 1 [ cs .MA] 8 Jul 2026  \nMulti-Agent Robotic Control with Onboard Vision-Language Models  \nKajetan Rachwał 1 ,2[0000−0003−1524−7877], Maciej Majek 1[0009−0009−9541−8461], Bartłomiej Boczek 1[0009−0006−9097−9655], Jakub Matejczyk 1[0009−0007−4835−1829], Dominik Matejkowski 1[0009−0008−6803−8001], Adam Dąbrowski 1[0000−0002−2130−0577], Tim Seyde3[0000−0001−9592−4465], Alexander Amini3[0000−0002−9673−1267], and Maria Ganzha2[0000−0001−7714−4844]  \n1 Robotec.AI, Warsaw, Poland  \n{[name.surname}@robotec.ai](name.surname}@robotec.ai)  \n2 Faculty of Mathematics and Information Science, Warsaw University of Technology,  \nWarsaw, Poland  \n[maria.ganzha@pw.edu.pl](maria.ganzha@pw.edu.pl)  \n3 Liquid.AI, Cambridge, Massachusetts, USA  \n{[name}@liquid.ai](name}@liquid.ai)  \nAbstract. Vision Language Models (VLMs) and Vision Language Action (VLA) models have shown promise in robotic control. Yet, they face significant challenges regarding explainability, generalization, and compute requirements. This paper presents a Multi-Agent System (MAS) architecture that addresses these limitations by deploying specialized agents on onboard hardware – eliminating dependence on external compute. The system controls a multi-purpose autonomous mobile manipulator in a simulated industrial warehouse, fulfilling five task categories:  \nsafety inspection, warehouse maintenance, warehouse search, package quality verification, and responding to human requests. Compact VLMs (3-20B parameters) are used throughout, with fine-tuning applied to improve package inspection accuracy. A novel “Megamind” orchestration agent mitigates context retention issues inherent to long-horizon planning with smaller models. The system was validated in a hardware-inthe-loop simulation using an AMD Ryzen ™ AI mini PC. Results demonstrate that a fully onboard MAS architecture is a viable, cost-efficient alternative to cloud-dependent deployments, with strong potential for real-world transfer. The simulation environment has been released as open source under the Apache 2 .0 licence.  \nKeywords: Multi-Agent Systems · Vision Language Models · Mobile Manipulation · Warehouse Robotics · Edge AI  \n1 Introduction  \nVision Language Models (VLMs) and Vision Language Action (VLA) models have been successfully utilized in the field of robotics in control tasks. However, these solutions have faced certain drawbacks. VLM based solutions typically  \n2 K. Rachwał et al.  \nrequire network connectivity and significant computing power in order to maintain an acceptable standard of operation. VLAs sometimes face similar issues. Their primary problem, however, is a general lack of explainability. This raises safety and ethical concerns. Some solutions to address these concerns exist, such as utilizing interpretability techniques on hidden layers or embodied chain-ofthought to intervene in steering [2] . Other authors propose utilizing constraint functions [6], or constraining deployment to “safe” robots [4] . However, all of these methods suffer from problems with generalization capabilities on out-ofdistribution data.  \nThe goal of this article is to demonstrate how these problems could be addressed utilizing Multi-Agent-System-based architecture (MAS) . To demonstrate this, a multipurpose autonomous mobile manipulator has been deployed. It’s controlled by a MAS (Fig. 1) running fully on the robot’s hardware. The system is capable of real-time response and operation in dynamic industrial environments. It has been tested in a hardware-in-the-loop simulation of a warehouse (Fig. 2) . The simulated robot platform uses RB-KAIROS+ coupled with an AMD Ryzen ™ AI mini PC for inference of the VLMs employed by the system’s agents. The VLM deployed on the robot’s hardware was LFM2-VL-3B [1] . Utilization of hardwarein-the-loop allows for feasibility assessment of using small locally hosted VLMs for robot control.  \nFig. 1. Deployed multi-agent architecture for vision-guided robotic inspection a","cbCairPwAussmGlK","https://ap.wps.com/l/cbCairPwAussmGlK","pdf",1218552,1,6,"English","en",105,"# Abstract\n# Introduction\n## Objectives\n# Main Purpose","[{\"question\":\"Why do vision-language approaches face limitations in robotic control?\",\"answer\":\"They often require network connectivity and significant compute to maintain acceptable performance, and many solutions lack explainability, raising safety and ethical concerns. Generalization to out-of-distribution data can also be weak.\"},{\"question\":\"What architecture is proposed to address explainability, generalization, and compute challenges?\",\"answer\":\"The paper presents a Multi-Agent System (MAS) where specialized agents run fully on onboard hardware. This removes dependence on external compute while using compact VLMs and a dedicated orchestration mechanism.\"},{\"question\":\"How does the system handle long-horizon planning with smaller vision-language models?\",\"answer\":\"It introduces a “Megamind” orchestration agent to mitigate context retention issues that arise when using smaller models for long-horizon planning.\"}]",1784186132,15,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"multi-agent-robotic-control-with-onboard-vision-language-models","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/multi-agent-robotic-control-with-onboard-vision-language-models/83236/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do vision-language approaches face limitations in robotic control?","Question",{"text":75,"@type":76},"They often require network connectivity and significant compute to maintain acceptable performance, and many solutions lack explainability, raising safety and ethical concerns. Generalization to out-of-distribution data can also be weak.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What architecture is proposed to address explainability, generalization, and compute challenges?",{"text":80,"@type":76},"The paper presents a Multi-Agent System (MAS) where specialized agents run fully on onboard hardware. This removes dependence on external compute while using compact VLMs and a dedicated orchestration mechanism.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the system handle long-horizon planning with smaller vision-language models?",{"text":84,"@type":76},"It introduces a “Megamind” orchestration agent to mitigate context retention issues that arise when using smaller models for long-horizon planning.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]