[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86214-en":3,"doc-seo-86214-105":29,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":20,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},86214,1374391974564,"Clementine","https://ap-avatar.wpscdn.com/avatar/14000253aa45c000a9e?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779874745381141002",8,"Research & Report","OpsMem Dual-Memory Reasoning with Cross-Memory Resonance for Failure Diagnosis","Failure diagnosis in modern software systems relies on iterative evidence acquisition and hypothesis reasoning grounded in operational experience. Existing LLM-based approaches improve diagnosis via agentic reasoning or knowledge augmentation, yet often fail to coordinate the evolving diagnostic state with experience across iterations. OpsMem introduces a dual-memory framework: short-term state memory and long-term reusable operational experience. Cross-memory resonance activates state-relevant experience to condition multi-agent diagnosis, then consolidates resolved incident experience back into long-term memory. Experiments on a real Huawei microservice dataset show significant gains in Match and Relevant.","OpsMem: Dual-Memory Reasoning with Cross-Memory Resonance for Failure Diagnosis  \nYongqian Sun†, Rongchen Gao†, Yu Luo†, Wenwei Gu†, Shenglin Zhang†∗ Qingyi Guo§ , Qiuai Fu‡, Yaoliang Wu‡, Dan Pei§†Nankai University § Tsinghua University ‡Huawei Technologies Co., Ltd.  \narXiv :2607 . 1 1357v 1 [ cs .AI] 13 Jul 2026  \nAbstract—Failure diagnosis in modern software systems requires iterative evidence acquisition and hypothesis reasoning guided by operational experience. Existing LLM-based methods improve diagnosis through agentic reasoning or knowledge augmentation, but they often lack a mechanism to coordinate the evolving diagnostic state with operational experience during iterative diagnosis. We propose OpsMem, a dual-memory framework that maintains a short-term memory for the current diagnostic state and a long-term memory for reusable operational experience. OpsMem uses cross-memory resonance to activate staterelevant long-term memory, conditions multi-agent diagnosis on the short-term and activated long-term memories, and consolidates reusable experience from solved incidents back into longterm memory. Experiments on a real-world Huawei microservice failure diagnosis dataset show that OpsMem outperforms representative agentic-reasoning and knowledge-augmented baselines, improving Match and Relevant by up to 46.88% and 18.39% over the strongest baseline, respectively.  \nIndex Terms—failure diagnosis, large language models, multiagent systems, agent memory  \nI. INTRODUCTION  \nFailure diagnosis is critical for maintaining the reliability of modern software systems, as engineers must quickly identify root causes and restore affected services when failures occur [1] . Traditionally, this task has relied heavily on manual inspection and expert experience, which is costly, time-consuming, and difficult to scale in today’s large-scale distributed environments [2] . To automate this process, prior studies have explored data-driven methods based on machine learning and deep learning [3], [4] . Although these methods have shown promise in specific scenarios, they often depend on stable data distributions, predefined failure patterns, and sufficient labeled data [5] . As a result, they suffer from limited generalizability and interpretability, making them difficult to apply reliably in real-world operations [6] .  \nRecent advances in large language models (LLMs) have demonstrated remarkable capabilities in language understanding, knowledge utilization, and complex reasoning [7], [8] . With these capabilities, LLMs can process diverse operational information and perform diagnostic reasoning, providing new opportunities for improving failure diagnosis [9]–[12] . In realworld operations, engineers diagnose failures through iterative evidence collection, observation analysis, and hypothesis refinement guided by operational experience [13] .  \n∗ Shenglin Zhang is the corresponding author.  \nTo support such iterative diagnosis, ReAct [14] enables LLMs to interleave reasoning and actions, forming a feedback loop between evidence acquisition and hypothesis refinement. However, in traditional ReAct-style diagnosis, the evolving diagnostic trajectory is often accumulated as a linear sequence of thoughts, actions, and observations in the context window. Asthe trajectory grows longer, such linear context becomes less reliable, making long-horizon diagnosis unstable [15], [16] . GoS [13] highlights typical long-horizon failures, including evidence fabrication, context drift, and failed backtracking, and mitigates them by organizing evidence and hypotheses into a structured belief state. Nevertheless, reliable diagnosis still requires operational experience for guidance.  \nAlthough LLMs acquire broad parametric knowledge during training, such knowledge is often insufficient for real-world failure diagnosis, which relies heavily on system-specific operational experience [17] . Existing methods usually incorporate such experience through retrieval, wh","cbCaij6WFGP6Ibxc","https://ap.wps.com/l/cbCaij6WFGP6Ibxc","pdf",1590853,6,1,"English","en",105,"# Introduction\n## Background and limitations of existing methods\n## LLM-based diagnosis and iterative evidence reasoning\n## Agentic reasoning vs. knowledge augmentation\n## Challenges for dual-memory design","[{\"question\":\"What evidence supports OpsMem’s effectiveness?\",\"answer\":\"Experiments on a real Huawei microservice failure diagnosis dataset show OpsMem outperforms representative agentic-reasoning and knowledge-augmented baselines, improving Match and Relevant by up to 46.88% and 18.39%.\"}]",1784209519,15,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":27},"opsmem-dual-memory-reasoning-with-cross-memory-resonance-for-failure-diagnosis","",{"@graph":35,"@context":77},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/opsmem-dual-memory-reasoning-with-cross-memory-resonance-for-failure-diagnosis/86214/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"What evidence supports OpsMem’s effectiveness?","Question",{"text":75,"@type":76},"Experiments on a real Huawei microservice failure diagnosis dataset show OpsMem outperforms representative agentic-reasoning and knowledge-augmented baselines, improving Match and Relevant by up to 46.88% and 18.39%.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":84},[85,89,93,97,102,106,111,114,119,122,126],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":45,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":45,"category_name":103,"show_sort_weight":104,"slug":105},"Technology",50,"technology",{"id":107,"doc_module":4,"doc_module_name":45,"category_name":108,"show_sort_weight":109,"slug":110},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":112,"slug":113},30,"research-report",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},9,"Religion & Spirituality",20,"religion-spirituality",{"id":117,"doc_module":4,"doc_module_name":45,"category_name":120,"show_sort_weight":117,"slug":121},"World Cup","world-cup",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":123,"slug":125},10,"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":98,"slug":129},19,"General","general"]