[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86546-en":3,"doc-seo-86546-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86546,962075114101,"Seraphina","https://ap-avatar.wpscdn.com/avatar/e000253a75eb197efd?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780044092746381165",8,"Research & Report","Recovery Control in Replicated Systems through Autonomous Multiagent Rollout","Recovery control is studied for replicated computing systems that provide service through multiple replicas working together. Redundancy allows continued operation if failed replicas are recovered faster than new failures arise. The timing decision for initiating recovery of selected replicas is modeled as a multiagent, partially observable Markov decision process (POMDP). A multiagent rollout approximation is used with precomputed signaling information to reduce coordination needs and enable parallel computation. Experiments demonstrate scalability to up to 70 replicas and lower recovery costs than common practical policies.","Recovery Control in Replicated Systems through Autonomous Multiagent Rollout  \nKim Hammar and Yuchao Li  \narXiv :2607 . 11187v1 [ ee ss . SY] 13 Jul 2026  \nAbstract—We study recovery control in replicated computing systems. Such systems consist of replicas that collectively provide a service to a client population. This redundancy enables the system to withstand failures provided that failed replicas are recovered faster than new failures occur. We show that the problem of deciding when to initiate recovery of selected replicascan be formulated as a partially observable Markov decision problem (POMDP) with a multiagent structure. We exploit this structure to apply a multiagent rollout method for approximating optimal control policies. Our method uses precomputed signaling information that reduces the need for replica coordination and facilitates parallel computations. Experiments show that our method scales to systems with up to 70 replicas and reduces costs compared to the recovery policies currently used in practice.  \nIndex Terms—Reinforcement learning, multiagent, recovery.  \nI. INTRODUCTION  \nAS our reliance on on-line services grows, there is an in  \ncreasing demand for reliable systems that provide correct service without disruption. Research on fault-tolerant systems has almost a century-long history, with the seminal work by von Neumann [1], and Moore and Shannon [2] in 1956 . The early work focused on tolerance against hardware failures. Since then, the field has broadened to include tolerance against software bugs, operator mistakes, and cyberattacks; see e.g., Lamport et al. [3] and Avizienis [4] . The common approach to building a fault-tolerant service is redundancy, whereby the service is provided by a set of service replicas. Through such redundancy, failed replicas can be substituted by healthy replicas as long as they can coordinate their service responses. This coordination problem is known as consensus.1  \nGiven a suitable consensus protocol, a distributed system with N service replicas can tolerate up to f \u003C N faulty replicas, where a faulty replica can behave arbitrarily, i.e., Byzantine. For example, a faulty replica can stop responding to service requests (e.g., due to a power outage) or send incorrect responses to clients (e.g., due to a cyberattack) . Tolerating failures in this setting means that despite the presence of up to f faulty replicas, the system as a whole is still guaranteed to provide correct responses to service requests by clients.  \nThis research is supported by the Swedish Research Council; 2024-06436 .  \nK. Hammar is with the Department of Computing, Imperial College, London, United Kingdom, [k.hammar@imperial.ac.uk](k.hammar@imperial.ac.uk).  \nY. Li is with the Fulton School of Engineering, Arizona State University, Tempe, [AZ.](AZ. yuchaoli@asu.edu)[ yuchaoli@asu.edu](AZ. yuchaoli@asu.edu).  \nManuscript received Month Day, 2026 .  \n1Here we refer to the classical notion of consensus from distributed systems theory, where a finite set of nodes must eventually agree on a value despite failures; see e.g., Lamport et al. [3] . This is distinct from the notion of consensus studied in control theory, which focuses on states gradually converging to the same value over time; see e.g., Olfati-Saber et al. [5] .  \nClients  \nService requests Responses  \nFig. 1. Multiagent recovery control in a system with N replicas that provide a service to clients, e.g., a web service. Each replica i ∈ {1, . . . , N} has a state xi , which is 1 if the replica is faulty (e.g., due to a cyberattack or operator mistake); 0 otherwise. Replicas are coordinated through a consensus protocol that ensures correct service as long as at most f \u003C N replicas are faulty. Each replica i is equipped with an agent that monitors the replica via the observation zi and applies the recovery control ui.  \nWhile the system can only tolerate f replicas being faulty simultaneously, it can withstand any number of failures provided t","cbCairl9hcemeGMH","https://ap.wps.com/l/cbCairl9hcemeGMH","pdf",4708770,4,1,15,"English","en",105,"# Introduction\n# Recovery Control Problem Formulation\n## Partially Observable Markov Decision Process (POMDP)\n## Multiagent Structure and Rollout Method\n# Related Approaches and Limitations\n# Proposed Scalable Recovery Method\n## Cost Reduction and Service Disruption","[{\"question\":\"What problem does the document address in replicated computing systems?\",\"answer\":\"It addresses when to initiate recovery of selected replicas so that service remains correct while balancing recovery costs against service requirements.\"},{\"question\":\"How is the recovery decision modeled?\",\"answer\":\"The timing decision is formulated as a partially observable Markov decision problem (POMDP) with a multiagent structure, where each replica is controlled based on partial observations.\"},{\"question\":\"Why does the proposed multiagent rollout method help performance?\",\"answer\":\"It leverages the multiagent structure and uses precomputed signaling information to reduce replica coordination overhead and facilitate parallel computation, yielding lower costs in experiments.\"}]",1784212545,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"recovery-control-in-replicated-systems-through-autonomous-multiagent-rollout","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/recovery-control-in-replicated-systems-through-autonomous-multiagent-rollout/86546/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-28","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the document address in replicated computing systems?","Question",{"text":75,"@type":76},"It addresses when to initiate recovery of selected replicas so that service remains correct while balancing recovery costs against service requirements.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is the recovery decision modeled?",{"text":80,"@type":76},"The timing decision is formulated as a partially observable Markov decision problem (POMDP) with a multiagent structure, where each replica is controlled based on partial observations.",{"name":82,"@type":73,"acceptedAnswer":83},"Why does the proposed multiagent rollout method help performance?",{"text":84,"@type":76},"It leverages the multiagent structure and uses precomputed signaling information to reduce replica coordination overhead and facilitate parallel computation, yielding lower costs in experiments.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]