[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85494-en":3,"doc-seo-85494-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85494,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","Cost and Accuracy of Long-Term Memory in Distributed Multi-Agent Systems Based on Large Language Models","Long-term memory (LTM) is central to large language model (LLM)-based agents in the internet of agents, yet existing evaluations often emphasize token usage and latency while overlooking system-level cost and deployment constraints in distributed multi-agent systems. The work proposes an independent, reproducible testbed measuring accuracy, latency, CPU time, peak RAM, disk I/O, and network usage in a simulated cloud-edge environment. It compares mem0, Graphiti, and cognee against RAG and full-context baselines on the LoCoMo benchmark. Results show a clear accuracy–cost clustering and identify retrieval incompleteness and compression precision as key drivers under unconstrained and constrained networks.","© IEEE 2026. Manuscript accepted at IEEE COMPSAC 2026. Not for redistribution. Published version: [https://doi.org/10.1109/XXXXXX](https://doi.org/10.1109/XXXXXX)  \nCost and Accuracy of Long-Term Memory in Distributed Multi-Agent Systems Based on Large  \nLanguage Models  \nBenedict Wolff  \nSchool of Electrical Engineering and Computer Science KTH Royal Institute of Technology Stockholm, Sweden [bjpwolff@kth.se](bjpwolff@kth.se)[ ](bjpwolff@kth.se)[bjpw@live.de](bjpw@live.de)  \nJacopo Bennati  \nSchool of Electrical Engineering and Computer Science KTH Royal Institute of Technology Stockholm, Sweden [jbennati@kth.se](jbennati@kth.se)[ ](jbennati@kth.se)[jacobbista.bennati@gmail.com](jacobbista.bennati@gmail.com)  \narXiv :2601 .07978v4 [ cs .IR] 13 Jul 2026  \nAbstract—Long-term memory (LTM) is fundamental to large language model (LLM)-based agents in the emerging internet of agents (IoA), where distributed multi-agent systems (DMASs) span cloud and edge networks. Existing evaluations are typically published by framework providers and focus on token usage and latency, rarely accounting for system-level cost or deployment in DMAS. These gaps are addressed with an independent reproducible testbed that evaluates accuracy, latency, CPU time, peak RAM, disk I/O and network usage in a simulated cloud-edge environment. Three venture capital-funded frameworks spanning vector, graph, and hybrid architectures, namely mem0, Graphiti, and cognee, are compared alongside retrieval-augmented generation (RAG) and full-context baselines on the LoCoMo benchmark under unconstrained and constrained network scenarios. Two clusters emerge: mem0, RAG, and full-context reach 77% to 81% accuracy, while Graphiti and cognee reach only 55% to 56%, a gap driven by retrieval incompleteness rather than reasoning failure. The RAG baseline matches the upper cluster at 8.4 times lower total cost of ownership (TCO) than mem0, and both are the only non-dominated backends on the Pareto frontier. Latency and bandwidth constraints as well as jitter leave retrieval quality unchanged for every backend, while vector-based LTM incurs a modest latency penalty of 4% to 5% under edge-cloud constraints. Compression precision rather than context volume determines LTM accuracy, as full-context forwarding underperforms mem0 despite supplying the entire conversation for each question.  \nIndex Terms—long-term memory, multi-agent system, internet of agents, retrieval-augmented generation, knowledge graph  \nOpen-source research: the testbed, results, and notebooks for analysis and evaluation are available at: [https://github.com/](https://github.com/)[ ](https://github.com/)[wolffbe/dmas-memory](wolffbe/dmas-memory)  \nI. INTRODUCTION  \nLarge language model (LLM)-based agents are increasingly capable of decomposing goals into executable steps using incontext learning, integrating external tools and data sources as well as persisting thoughts using long-term memory (LTM) [1] . LTM refers to the ability of LLM-based agents to store, organize, and retrieve information across interactions, enabling persistent context, behavioral adaptation, and personalization beyond single-turn reasoning [2] . Retrieval-augmented  \ngeneration (RAG) is foundational to LTM by embedding overlapping chunks of data into high-dimensional vector representations [3] . However, this static approach struggles with dynamic knowledge integration, introduces noise, and overlooks structural relationships. LTM extends this paradigm into dynamic, agent-specific memory that accumulates overtime, organized through textual, vector-based, graph-based and hybrid architectures, and supported by operations for data and knowledge indexing, retrieval, updating, and consolidation.  \nTo evaluate these capabilities, LTM benchmarks such as LoCoMo and LongMemEval have been proposed, focusing on tasks including multi-session reasoning, information retrieval, temporal consistency, and knowledge updates [4], [5] .  \nDespite advances in L","cbCaivvH251BUj2S","https://ap.wps.com/l/cbCaivvH251BUj2S","pdf",341335,2,1,11,"English","en",105,"# Introduction\n# Proposed Evaluation Testbed and Metrics\n# Compared LTM Frameworks and Baselines\n# Results and Analysis under Network Constraints\n# Open-Source Resources","[{\"question\":\"What problem does the paper address about LTM evaluation in distributed multi-agent systems?\",\"answer\":\"Existing evaluations often focus on token and latency metrics published by framework providers, and rarely measure system-level cost or deployment behavior in cloud-edge distributed settings.\"},{\"question\":\"What does the proposed testbed measure besides accuracy?\",\"answer\":\"It evaluates accuracy, latency, CPU time, peak RAM, disk I/O, and network usage using an independent and reproducible setup.\"},{\"question\":\"What explains the performance gap between the two accuracy clusters?\",\"answer\":\"The paper attributes the gap mainly to retrieval incompleteness rather than reasoning failure.\"}]",1784204009,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"cost-and-accuracy-of-long-term-memory-in-distributed-multi-agent-systems-based-on-large-language-models","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/cost-and-accuracy-of-long-term-memory-in-distributed-multi-agent-systems-based-on-large-language-models/85494/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address about LTM evaluation in distributed multi-agent systems?","Question",{"text":75,"@type":76},"Existing evaluations often focus on token and latency metrics published by framework providers, and rarely measure system-level cost or deployment behavior in cloud-edge distributed settings.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does the proposed testbed measure besides accuracy?",{"text":80,"@type":76},"It evaluates accuracy, latency, CPU time, peak RAM, disk I/O, and network usage using an independent and reproducible setup.",{"name":82,"@type":73,"acceptedAnswer":83},"What explains the performance gap between the two accuracy clusters?",{"text":84,"@type":76},"The paper attributes the gap mainly to retrieval incompleteness rather than reasoning failure.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]