[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84653-en":3,"doc-seo-84653-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84653,3848291630094,"Emma Wilson","https://eur-avatar.wpscdn.com/davatar_085a072bc5b1113ac321206ff7593b45",8,"Research & Report","Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots","Embodied.cpp presents a portable C++ inference runtime for embodied AI models, addressing the fragmented deployment landscape across model-specific Python stacks, robot-side glue code, and backend assumptions. The runtime targets an embodied inference contract with modular multi-rate execution inside closed-loop control, latency-first batch-1 operation on heterogeneous edge hardware, and extensible operator plus embodied I/O beyond fixed token interfaces. Evaluations on HY-VLA and pi0.5 achieve 100.0% and 91.0% task success, while a WAM benchmark reduces block memory from 312.2 MiB to 88.1 MiB.","arXiv :2607 .0250 1v2 [ cs .RO] 3 Jul 2026  \n2026-07-01  \nEmbodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots  \nLing Xu 1 Chuyu Han2 Borui Li 1,† Hao Wu2,† Shiqi Jiang3 Ting Cao4 Chuanyou Li 1 Sheng Zhong2 Shuai Wang 1  \n1  Southeast University 2  Nanjing University 3  Microsoft Research  \n4  Institute for AI Industry Research (AIR), Tsinghua University †Project Leader  \nEmbodied AI models now span vision-language-action (VLA) models and world-action models (WAMs), but practical deployment remains fragmented across model-specific Python stacks, backend assumptions, and robot-side glue code, especially on heterogeneous edge devices. Existing inference runtimes are designed mainly for request-response serving and therefore do not satisfy the runtime contract of embodied deployment: multi-rate execution inside closed-loop control, latency-first batch-1 inference on heterogeneous hardware, and extensible embodied interfaces beyond fixed token I/O. We present Embodied .cpp, a portable C++ inference runtime for embodied models. Based on an architectural analysis of representative VLA models and WAMs, Embodied .cpp captures a shared execution path and organizes it into five layers: input adapters, sequence builders, backbone execution, head plugins, and deployment adapters. The runtime provides modular multi-rate execution, latency-first fused inference, and extensible operator and I/O support, enabling deployment across heterogeneous devices, robots, and simulators through one backend abstraction. We evaluate Embodied .cpp on two VLA models, HY-VLA and pi0.5, and on a preliminary WAM benchmark using a LingBot-VA Transformer block. The VLA deployments achieve successful closed-loop execution with 100.0% and 91.0% task success rates, respectively. The WAM benchmark reduces block memory from 312.2 MiB to 88.1 MiB. These results show that Embodied .cpp improves deployment efficiency while preserving high accuracy across diverse embodied model architectures.  \nProject Link: [https://github.com/SEU-PAISys/Embodied.cpp](https://github.com/SEU-PAISys/Embodied.cpp)  \n1. Introduction  \nAcademia and industry have already produced a rapidly growing set of embodied models. Recent systems now span vision-language-action (VLA) models and world-action models (WAMs), showing substantial progress in model architecture and capability [1, 2, 3, 4, 5, 6, 7] . At the same time, the surrounding ecosystem for data, training, and simulation has also improved substantially through projects such as LeRobot, Open X-Embodiment, ManiSkill, LIBERO, and Isaac Sim [8, 9, 10, 11, 12] . But practical impact does not end at model construction. To become reliable robot-side systems, these models must run on heterogeneous and resource-constrained devices, from Jetson and RK-based boards to x86 edge boxes and workstation-class robots, without rebuilding the software stack for each new model family.  \nAchieving this requires a unified inference runtime. Existing inference runtimes, however, do not yet provide such support. General LLM or VLM runtimes are designed for request-response serving, relatively  \nuniform token interfaces, and throughput-oriented optimization [13, 14, 15, 16] . Embodied inference, by contrast, lives inside a closed-loop interaction process with robot-and simulator-side dependencies. As a result, even a strong checkpoint still has to be stitched into Python research code, backend-specific inference paths, handwritten sensor wrappers, and platform-specific control logic before it can act on a robot. As embodied architectures diversify, this integration burden only grows, making a portable inference runtime for embodied AI models increasingly necessary.  \nThe challenge arises because embodied deployment fundamentally changes the runtime contract. Compared with conventional LLM or VLM serving, an embodied inference runtime must satisfy three key requirements. (1) Multi-rate execution. Embodied inference is not a si","cbCaihZLwyxf1t2Z","https://ap.wps.com/l/cbCaihZLwyxf1t2Z","pdf",2315768,1,12,"English","en",105,"# Introduction\n## Key challenges in embodied inference deployment\n## Runtime requirements: multi-rate, latency-first control, extensible interfaces\n## Embodied.cpp design and contributions\n## Experimental evaluation highlights","[{\"question\":\"What problem does Embodied.cpp target in embodied AI deployment?\",\"answer\":\"Embodied.cpp targets the fragmentation caused by model-specific Python stacks, differing backend assumptions, and robot-side integration glue that prevents models from running reliably on heterogeneous devices and robots.\"},{\"question\":\"What three requirements define the runtime contract for embodied inference?\",\"answer\":\"It requires modular multi-rate execution in closed-loop control, latency-first batch-1 operation with low jitter on heterogeneous edge hardware, and extensible embodied interfaces supporting custom operators and non-fixed I/O formats.\"},{\"question\":\"What results are reported for VLA and WAM evaluations?\",\"answer\":\"HY-VLA and pi0.5 achieve 100.0% and 91.0% task success rates in closed-loop execution, while a preliminary WAM benchmark reduces block memory from 312.2 MiB to 88.1 MiB.\"}]",1784197504,30,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"embodiedcpp-a-portable-inference-runtime-of-embodied-ai-models-on-heterogeneous-robots","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/embodiedcpp-a-portable-inference-runtime-of-embodied-ai-models-on-heterogeneous-robots/84653/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does Embodied.cpp target in embodied AI deployment?","Question",{"text":75,"@type":76},"Embodied.cpp targets the fragmentation caused by model-specific Python stacks, differing backend assumptions, and robot-side integration glue that prevents models from running reliably on heterogeneous devices and robots.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What three requirements define the runtime contract for embodied inference?",{"text":80,"@type":76},"It requires modular multi-rate execution in closed-loop control, latency-first batch-1 operation with low jitter on heterogeneous edge hardware, and extensible embodied interfaces supporting custom operators and non-fixed I/O formats.",{"name":82,"@type":73,"acceptedAnswer":83},"What results are reported for VLA and WAM evaluations?",{"text":84,"@type":76},"HY-VLA and pi0.5 achieve 100.0% and 91.0% task success rates in closed-loop execution, while a preliminary WAM benchmark reduces block memory from 312.2 MiB to 88.1 MiB.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":28,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]