[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82842-en":3,"doc-seo-82842-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82842,5909877438554,"Maeve","https://ap-avatar.wpscdn.com/avatar/5600025385ad2bf12a7?_k=1778553567797529272",8,"Research & Report","Can LLMs Really Recover Microservice Failures? A Recovery-Aware Evaluation of Diagnosis-to-Action Reasoning","Large language models (LLMs) are used to interpret operational evidence and support incident response in cloud-native microservice systems, yet recovery-oriented scenarios require more than root-cause identification. Operators must convert diagnosis into a concrete recovery action, select an admissible target, and verify service health restoration. R2Act is proposed as a recovery-action evaluation framework, including an incident schema, actionspace modeling, recovery-validity metrics, an offline evaluator, and a live-replay protocol.","Can LLMs Really Recover Microservice Failures? A Recovery-Aware Evaluation of Diagnosis-to-Action Reasoning  \narXiv :2607 .04623v 1 [ cs . SE] 6 Jul 2026  \nJiaxing Qi∗ , Zhongzhi Luan∗ , Hongyu Zhang†, Shaohan Huang‡  \nCarol Fung§ , Yongxin Tong∗ , Hailong Yang∗ , and Depei Qian∗  \n∗ Beihang University  \n[jiaxingqi@buaa.edu.cn](jiaxingqi@buaa.edu.cn), [luan.zhongzhi@buaa.edu.cn](luan.zhongzhi@buaa.edu.cn), [yxtong@buaa.edu.cn](yxtong@buaa.edu.cn), [hailongyang@buaa.edu.cn](hailongyang@buaa.edu.cn), [depeiq@buaa.edu.cn](depeiq@buaa.edu.cn)  \n†Chongqing University  \n[hyzhang@cqu.edu.cn](hyzhang@cqu.edu.cn)  \n‡Microsoft Research Asia  \n[shaohanh@microsoft.com](shaohanh@microsoft.com)  \n§ Concordia University  \ncarol.fung@concordia.ca  \nAbstract—Large language models (LLMs) are increasingly used to interpret operational evidence and assist incident response in cloud-native microservice systems. However, recoveryoriented use cases require more than identifying a root cause. After observing symptoms and diagnosing a fault, an operator or agent must translate the diagnosis into a concrete recovery action, apply it to an admissible target, and verify that service health has been restored. Existing RCA and log-analysis evaluations are well-suited to diagnosis, but they do not characterize this subsequent action decision. This paper presents R2Act, a recovery-action evaluation framework for post-diagnosis incident response. R2Act defines an incident schema, quality gate, actionspace representation, recovery-validity metrics, offline evaluator, and live-replay protocol. We instantiate the framework as a benchmark dataset of 302 quality-audited Kubernetes incidents from Online Boutique. Each incident provides synchronized multi-modal observations, root-cause labels, an incident-specific action space, and annotated valid and invalid recovery plans. We evaluate heuristic, supervised, RCA-oriented, deep log, and LLMbased methods. The strongest RAG-based LLMs reach 91.4%– 99.7% root-cause service accuracy, yet their recovery validity remains only 36.8%–60.3% . Even when both the root-cause service and fault type are correct, recovery-oriented methods still choose invalid actions for 39.5%–62.0% of correctly diagnosed incidents. In validity-gated live replay, 146 of 302 Qwen-RAG predictions are both offline-valid and recovered in live execution. Overall, this work reveals that many recovery failures arise not from missing diagnostic knowledge, but from the difficulty of translating diagnostic evidence into valid recovery actions and admissible targets. This work provides a reproducible, simplified starting point for research and evaluation.  \nIndex Terms—microservices, fault diagnosis, recovery planning, recovery-aware evaluation, LLMs  \nI. INTRODUCTION  \nAI-assisted operations are advancing rapidly along several axes. LLM-based operational systems can summarize logs, retrieve similar incidents, localize likely root causes, and draft mitigation steps [1], [2] . At the same time, cloud-native microservice systems have become harder to recover because  \nFig. 1. R2Act adds a recovery-action evaluation layer after diagnosis, checking whether methods can select a valid operation and target under incident-specific constraints.  \nindependent deployment, elastic scaling, and dense service dependencies create more ways for failures to propagate [3],[4], [5] . A user-visible incident may leave evidence scattered across logs, Kubernetes events, metrics, and resource state. These trends change incident response from a diagnosis-only task into an action-oriented decision problem: an automated method must not only explain what failed, but also decide which recovery action should be taken and where that action should be applied.  \nYet there have been relatively few efforts in this recovery process. Real microservice recovery begins with partial, noisy, and multi-modal operational evidence, while the required outcome is a concrete decision that may modify a l","cbCaicHfyuAigZPH","https://ap.wps.com/l/cbCaicHfyuAigZPH","pdf",7019869,1,12,"English","en",105,"# Introduction\n## Background and problem motivation\n## Related work and identified gaps","[{\"question\":\"What problem does this paper focus on for LLMs in microservice incident response?\",\"answer\":\"It focuses on whether LLM-based methods can translate post-diagnosis information into valid recovery actions, including selecting an admissible target and parameters, not just identifying root causes.\"},{\"question\":\"What is R2Act and what components does it include?\",\"answer\":\"R2Act is a recovery-action evaluation framework that defines an incident schema, represents action spaces, introduces recovery-validity metrics, and supports both offline evaluation and a live-replay protocol.\"},{\"question\":\"What do the experiments show about diagnostic accuracy versus recovery validity?\",\"answer\":\"Even when root-cause service and fault type are correct, recovery-oriented methods still select invalid actions frequently, indicating that translating diagnostic evidence into valid recovery decisions is the main challenge.\"}]",1784183363,30,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"can-llms-really-recover-microservice-failures-a-recovery-aware-evaluation-of-diagnosis-to-action-reasoning","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/can-llms-really-recover-microservice-failures-a-recovery-aware-evaluation-of-diagnosis-to-action-reasoning/82842/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does this paper focus on for LLMs in microservice incident response?","Question",{"text":75,"@type":76},"It focuses on whether LLM-based methods can translate post-diagnosis information into valid recovery actions, including selecting an admissible target and parameters, not just identifying root causes.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is R2Act and what components does it include?",{"text":80,"@type":76},"R2Act is a recovery-action evaluation framework that defines an incident schema, represents action spaces, introduces recovery-validity metrics, and supports both offline evaluation and a live-replay protocol.",{"name":82,"@type":73,"acceptedAnswer":83},"What do the experiments show about diagnostic accuracy versus recovery validity?",{"text":84,"@type":76},"Even when root-cause service and fault type are correct, recovery-oriented methods still select invalid actions frequently, indicating that translating diagnostic evidence into valid recovery decisions is the main challenge.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":28,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]