[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85210-en":3,"doc-seo-85210-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85210,962075114101,"Seraphina","https://ap-avatar.wpscdn.com/avatar/e000253a75eb197efd?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780044092746381165",8,"Research & Report","ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory","ABot-AgentOS is a general robotic Agent Operating System built above low-level controllers to support long-horizon embodied agents. The framework provides deliberative, scene-conditioned planning, context-isolated skill execution, multi-stage verification, multi-modal memory, and edge-cloud collaboration. It also introduces EmbodiedWorldBench, an executable benchmark with 16 indoor, outdoor, and hybrid scenes, four difficulty levels, and 200+ tasks for navigation, object search, NPC dialogue, and dynamic events. Using Universal Multi-modal Graph Memory and a failure-driven self-evolution loop, ABot-AgentOS improves task success and goal completion while preventing cross-split memory leakage.","arXiv :2607 . 10350v 1 [ cs .AI] 11 Jul 2026  \nABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory  \nAMAP CV Lab  \nSee Contributions section for a full author list.  \nAbstract  \nRecent VLM and VLA systems have improved robotic perception and action prediction, yet long-horizon embodied agents still require a general runtime layer for reasoning, memory, tool use, verification, and cross-embodiment execution. We present ABot-AgentOS, a general robotic Agent Operating System that sits above low-level controllers and provides a deliberative agent layer for scene-conditioned planning, context-isolated skill execution, multi-stage verification, multi-modal memory, and edge-cloud collaboration. To evaluate such systems, we introduce EmbodiedWorldBench, an executable benchmark with 16 indoor, outdoor, and hybrid scenes, four difficulty levels, and over 200 tasks involving navigation, object search, NPC dialogue, dynamic events, and trace-grounded scoring.  \nABot-AgentOS further introduces Universal Multi-modal Graph Memory, a persistent sourcegrounded substrate that converts dialogue, visual observations, spatial context, temporal relations, and task traces into typed nodes and edges. A failure-driven self-evolution loop converts diagnosed memory failures into gated runtime evo-assets that are promoted only to later evaluation splits, preventing current-split ground-truth leakage while enabling continual improvement. On an initial EmbodiedWorldBench subset, ABot-AgentOS improves over a single-controller baseline in both task success and goal completion. Across memory benchmarks, ABot-AgentOS Static achieves 87.5 on LoCoMo, 59.9 on OpenEQA EM-EQA, 88.6 on Mem-Gallery, and 76.5 Acc@All on NExT-QA; self-evolution further improves LoCoMo to 88.7, OpenEQA to 60.4, and Mem-Gallery to 89.0 . These results suggest that a general Agent OS layer can improve long-horizon embodied execution while providing persistent, auditable memory for continual interaction.  \nDate: July 10, 2026  \nProject Page: [https://amap-cvlab.github.io/ABot-AgentOS](https://amap-cvlab.github.io/ABot-AgentOS)  \nContents  \n1 Introduction ................................................ 4  \n2 Agent Framework ............................................ 5  \n2.1 Architecture Overview ........................................ 5  \n2.2 Agent Harness ............................................ 6  \n2.2.1 Scene-Conditioned Task Planning .............................. 7  \n2.2.2 Skill Runner for Procedural Execution ........................... 8  \n2.2.3 Multi-Stage Verification ................................... 8  \n2.2.4 Edge-Cloud Collaborative Routing ............................. 8  \n2.3 Memory System ........................................... 9  \n2.3.1 Graph Memory Representation ............................... 10  \n2.3.2 Multi-modal Memory Updating ............................... 10  \n2.3.3 Retrieval and Grounded Answering ............................. 11  \n2.3.4 Failure-Driven Lifelong Self-Evolution ........................... 11  \n2.3.5 Edge-Cloud Collaborative Memory Management ..................... 13  \n3 EmbodiedWorldBench ......................................... 14  \n3.1 Overview ............................................... 14  \n3.2 Benchmark Construction ....................................... 14  \n3.2.1 Environment and Semantic Map Annotation ....................... 14  \n3.2.2 Query Generation ...................................... 15  \n3.3 Evaluation Procedure and Metrics ................................. 16  \n3.3.1 Embodied Evaluation Workflow ............................... 16  \n3.3.2 Evaluation Metrics ...................................... 16  \n4 Model Training .............................................. 17  \n4.1 Overview ............................................... 17  \n4.2 Text-Based Environment and Task Construction ......................... 18  \n4.3 Teacher Trajectory Distillation and SFT .............................. 18  \n4.","cbCaiviAb0cVkSzf","https://ap.wps.com/l/cbCaiviAb0cVkSzf","pdf",32291582,3,1,40,"English","en",105,"# Introduction\n# Agent Framework\n## Architecture Overview\n## Agent Harness\n### Scene-Conditioned Task Planning\n### Skill Runner for Procedural Execution\n### Multi-Stage Verification\n### Edge-Cloud Collaborative Routing\n## Memory System\n### Graph Memory Representation\n### Multi-modal Memory Updating\n### Retrieval and Grounded Answering\n### Failure-Driven Lifelong Self-Evolution\n### Edge-Cloud Collaborative Memory Management\n# EmbodiedWorldBench\n## Overview\n## Benchmark Construction\n## Evaluation Procedure and Metrics\n# Model Training\n## Overview\n## Training Pipeline\n# Experiments\n## Agent Evaluation\n## Memory Evaluation\n# Conclusion","[{\"question\":\"What core capability does ABot-AgentOS add for long-horizon embodied agents?\",\"answer\":\"ABot-AgentOS adds a general agent operating system layer that supports reasoning, memory, tool/skill execution, multi-stage verification, and edge-cloud collaboration on top of low-level controllers.\"},{\"question\":\"How does EmbodiedWorldBench evaluate robotic agents?\",\"answer\":\"EmbodiedWorldBench provides 16 mixed indoor/outdoor scenes with four difficulty levels and 200+ tasks covering navigation, object search, NPC dialogue, and dynamic events, along with trace-grounded scoring.\"},{\"question\":\"What is Universal Multi-modal Graph Memory and how is lifelong improvement handled?\",\"answer\":\"Universal Multi-modal Graph Memory converts dialogue, visual observations, spatial/temporal context, and task traces into typed nodes and edges. A failure-driven self-evolution loop converts diagnosed memory failures into runtime assets that are promoted only to later evaluation splits to avoid ground-truth leakage.\"}]",1784201760,101,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"abot-agentos-a-general-robotic-agent-os-with-lifelong-multi-modal-memory","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/abot-agentos-a-general-robotic-agent-os-with-lifelong-multi-modal-memory/85210/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What core capability does ABot-AgentOS add for long-horizon embodied agents?","Question",{"text":75,"@type":76},"ABot-AgentOS adds a general agent operating system layer that supports reasoning, memory, tool/skill execution, multi-stage verification, and edge-cloud collaboration on top of low-level controllers.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does EmbodiedWorldBench evaluate robotic agents?",{"text":80,"@type":76},"EmbodiedWorldBench provides 16 mixed indoor/outdoor scenes with four difficulty levels and 200+ tasks covering navigation, object search, NPC dialogue, and dynamic events, along with trace-grounded scoring.",{"name":82,"@type":73,"acceptedAnswer":83},"What is Universal Multi-modal Graph Memory and how is lifelong improvement handled?",{"text":84,"@type":76},"Universal Multi-modal Graph Memory converts dialogue, visual observations, spatial/temporal context, and task traces into typed nodes and edges. A failure-driven self-evolution loop converts diagnosed memory failures into runtime assets that are promoted only to later evaluation splits to avoid ground-truth leakage.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":22,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]